AIMC Journal:
bioRxiv

Showing 451 to 460 of 4935 articles

A self-supervised DNA foundation model with collapse-resistant multimodal fusion

bioRxiv
Genomic foundation models pretrained on DNA sequence have achieved strong performance across a range of tasks, but sequence-only representations cannot fully capture regulatory information reflected by additional DNA-centric modalities. Existing mult...

Automating scientific annotations for open transcriptomic profiles via multi-stage agents

bioRxiv
Public transcriptomic repositories contain millions of samples, yet their large-scale reuse is hindered by heterogeneous and inconsistently reported metadata. In the Gene Expression Omnibus (GEO), key biological information is often distributed acros...

PerturbTrace: Evaluating Feedback Use by AI Co-Scientist Agents in Perturbation Discovery

bioRxiv
Recent advances in AI co-scientists have brought LLM agents into closed-loop experimental design. However, whether these agents use feedback from earlier rounds to revise subsequent experimental decisions remains unclear. We address this question wit...

Microbial bioprospecting for benzoxazolinate-like molecules: unleashing the potential of genome mining

bioRxiv
The benzoxazolinate moiety is a key functional group found in a few natural products (NPs), exhibiting diverse bioactivities, including antitumor, antibacterial, and cytotoxic activities. Despite their clinical importance, only a few bacterial strain...

bulk2scDiff: A Pseudobulk-Conditioned Diffusion Model for Bulk-to-Single-Cell RNASeq Generation

bioRxiv
Bulk RNA sequencing remains the predominant profiling strategy for large clinical cohorts, but it aggregates transcriptional signals across cell populations, thereby masking the underlying cellular heterogeneity. Inferring this heterogeneity from exi...

ETAP-CLF: an ESM3-based transformer attention framework for binary protein classification

bioRxiv
Binary protein classification supports diverse tasks in computational biology, including pathway-membership inference and sequence-based candidate prioritization. Protein language models generate information-rich residue-level representations, but do...

A Semantic + Neuronal Approach to Predict Pathogenic Variants in DNA Sequences

bioRxiv
In this work, we present a machine learning model for identifying pathogenic DNA variants. The model was learned from the analysis of normal and pathogenic sequences extracted from the ClinVar database (supported by NCBI). This analysis was based on ...

GlycoMeSH: linking glycan structures to biomedical context for systematic enrichment analysis

bioRxiv
Glycan identification has advanced, but glycan structures remain difficult to translate into reproducible biomedical context because reusable glycan-level annotations are sparse. We present GlycoMeSH, a resource that links glycans to Medical Subject ...

PandaDock: An Open-Source Molecular Docking Platform with Flexible-Ligand Search and Equivariant Neural Scoring

bioRxiv
We present PandaDock, an open-source molecular docking platform implementing flexible-ligand conformational search with analytic gradients, a precomputed affinity grid engine, specialized modules for induced-fit, metal-coordination and tethered docki...

Sparse Autoencoders Reveal Structural and Family-level Features in BiRNA-BERT

bioRxiv
Motivation: RNA language models learn representations that support structure and function prediction, but which biological concepts their hidden states encode remains unclear. Sparse autoencoders (SAEs) decompose hidden states into interpretable feat...