The Rancho Resource Center for omics and biomedical data

Peer-reviewed publications, blogs, whitepapers, webinars, events, and scientific insights from the team harmonizing biomedical data for modern R&D.

Browse by Topic

Browse by Resource Type

Filters

Posters Posters
Multi-Omics Single-Cell

Accelerating Biomedical Discovery with Public Data: Single-Cell Data Science Consortium Harnesses Collective Expertise to Deliver Even More High Value Harmonized Datasets in Year 3

The Single-Cell Data Science (SCDS) Consortium is a pre-competitive collaboration led by Rancho BioSciences (launched 2022) that harmonizes publicly available single-cell sequencing datasets to overcome challenges of inconsistent metadata standards and variable processing approaches.

Rancho BioSciences
Posters Posters
Multi-Omics Single-Cell

Rancho Biosciences Single-Cell Data Science Consortium Enables Accelerated Discovery With 100-Million AI & Analysis-Ready Single-Cells

The Rancho Single-Cell Data Science Consortium has delivered 100 million AI-ready single-cells across 1,000 analysis-ready datasets to 11 member organizations, featuring harmonized processing, comprehensive FAIR metadata, and advanced cell-typing to accelerate therapeutic discovery. The consortium demonstrates utility through construction of disease atlases (Dermatitis and IBD) that validate existing therapeutic targets (IL13, IL4R, IL31) and identify novel candidates (IL26, IL32), showcasing the power of integrated single-cell data for precision medicine and drug discovery.

Rancho BioSciences
Posters Posters
Genomics Data Harmonization

Aggregation and integration of UK Biobank phenotypes data for downstream analysis

Rancho BioSciences and Takeda addressed challenges in the UK Biobank's 502,600-subject dataset (3,390+ variable fields) by implementing ~100 harmonization rules to standardize inconsistent data formats, converting categorical values to numerical formats, and mapping diagnoses, procedures, and medications to uniform ontologies (ICD10, SNOMED CT, RXNORM, MeSH). This curation approach enables algorithmic derivation of enriched, scientifically relevant phenotypes and increases disease cohort sizes by combining multiple diagnostic codes, thereby enhancing statistical power for downstream genomic and biomarker analyses.

Rancho BioSciences
Posters Posters
Ontology Data Harmonization

Application of an AI-Powered Terminology Management Solution (TMS) in the Real-World Data (RWD) FAIRification process

Rancho Biosciences' Terminology Management Solution (TMS) combines AI-assisted semantic and phonetic (Fuzzy) algorithms to automate terminology mapping across 50+ biomedical ontologies, demonstrating superior performance in mapping real-world data terms with 99% recall and up to 74% correct matches—outperforming commercial tools in precision, efficiency, and coverage. When integrated into an OMOP-based RWD harmonization pipeline for >10,000 terms, TMS saved ~500 hours of manual curation time while achieving a 14 percentage point increase in correct mappings through combined AI and Fuzzy approaches, making it a critical tool for FAIRifying real-world data for drug discovery and regulatory submissions.

Rancho BioSciences
Posters Posters
Data Harmonization Bioinformatics

Assessing the use of telemetric devices in preclinical safety studies in non-human primates - data aggregation step

Rancho Biosciences and Genentech aggregated and harmonized data from 30 preclinical toxicology studies (343 Cynomolgus monkeys) comparing telemetry-instrumented versus non-instrumented animals across 174 clinical, biochemical, and physiological parameters to assess instrumentation effects. Analysis revealed only 5% of parameters differed significantly between cohorts (primarily hematology and clinical chemistry markers like globulin, albumin/globulin ratio, and liver enzymes), with effect sizes calculated by dividing mean differences by standard deviation, demonstrating that surgical telemetry instrumentation has minimal impact on most toxicology endpoints and provides comparable data to conventional monitoring methods.

Rancho BioSciences
Posters Posters
Data Harmonization Bioinformatics

Automated Extraction of Biotransformation Data from Metabolic Schemas Using Geometric Criteria and Large Language Models

Rancho Biosciences and Procter & Gamble Human Safety developed an automated pipeline to digitize biotransformation pathways from >2,000 metabolism reports, combining geometric criteria from ChemDraw XML parsing with large language models (Claude-3.5) to extract chemical structures, reactions, metadata (biospecimen, species, references), and biotransformation types. The pipeline achieved 98.5% success in reaction graph extraction and 98.9% accuracy in reagent/product identification—outperforming manual extraction (96.8% and 92.3% respectively)—enabling creation of an AI-ready metabolism database to reduce animal testing, improve metabolite prediction models beyond current pharmaceutical-focused tools, and support safety assessments across broader chemical domains.

Rancho BioSciences
Posters Posters
Data Harmonization Bioinformatics

A Comprehensive Dataset of Pharmacokinetic Parameters for Recommended Doses of Drugs: Enabling Drug Repositioning and Pharmacological Analysis

Rancho BioSciences and the National Center for Advancing Translational Sciences (NCATS) created a centralized dataset of major PK parameters (Cmax, AUC, T1/2, fraction unbound) for recommended doses of 637 approved and investigational drugs, integrating detailed administration metadata (frequency, route, food status), population characteristics, and molecular descriptors from FDA labels, Phase III trials, and scientific literature. Analysis revealed strong Cmax-AUC correlation modulated by half-life, negative associations between LogP and fraction unbound, and tightening PK distributions in more recently approved drugs (1980s vs 2010s, p<0.01), providing a critical resource to connect in vitro efficacy with in vivo dosing for drug repositioning and pharmacological analysis across 51 target classes including small molecules and peptides.

Rancho BioSciences
Posters Posters
Multi-Omics Proteomics

Bringing Together Multi-Omics and Real-World Data to Accelerate Insight Delivery for Biomarker Discovery & Drug Development

This poster (a Rancho Biosciences/Sapient Bioanalytics collaboration) describes building a hybrid data lakehouse aligned to the OMOP common data model to integrate multi-omics data (proteomics, metabolomics, genomics) with real-world clinical data, using automated harmonization and curation to enable large-scale biomarker analyses. In a featured use case, this infrastructure was applied to >20,000 human samples to identify an early diagnostic biomarker for metabolic dysfunction-associated steatohepatitis (MASH), demonstrating the approach's ability to accelerate biomarker discovery and drug development.

Rancho BioSciences
Posters Posters
Data Harmonization Bioinformatics

Building towards a computational infrastructure to aid in the interpretation of Bicycle toxin conjugate response profiles

Bicycle Therapeutics and Rancho BioSciences developed an integrated computational pipeline analyzing pharmacogenomic data (gCSI, TCGA) to identify factors influencing response to Bicycle Toxin Conjugates (BTCs like BT1718, BT5528, BT8009), discovering 181 genes associated with microtubule binding agent (MTBA) sensitivity, with low ABCB1 (drug efflux pump) expression strongly correlated with MTBA sensitivity across cancer indications. The pipeline integrates tumor antigen expression (MT1-MMP, EphA2, Nectin-4) with ABCB1 levels to predict BTC response profiles, demonstrating that bladder cancer (BLCA) exhibits favorable characteristics of low ABCB1 and high target antigen expression, supporting a multiplex biomarker approach to elucidate patient response in ongoing Phase I/II clinical trials.

Rancho BioSciences
Posters Posters
Data Harmonization Bioinformatics

Data deluge: challenges and solutions in implementing a precision medicine approach to create the TRACK-TBI information Commons

Rancho BioSciences and UCSF developed comprehensive data curation workflows and validation infrastructure for the TRACK-TBI study (3,264+ TBI patients across 19 U.S. trauma centers), harmonizing clinical data (15,000+ fields), biospecimen logs (90,000+ samples), and neuroimaging datasets to NIH Common Data Elements for submission to FITBIR. The team created R Shiny applications and standardized biospecimen templates with automated validation rules that identified and corrected thousands of data entry errors (mismatched sample IDs, timestamp inconsistencies, protocol omissions), transforming disparate multi-modal datasets (clinical, proteomic, genomic, imaging, outcomes) into analysis-ready resources to enable precision medicine approaches for TBI biomarker discovery and future clinical trials, supported by NINDS, DoD, Abbott Laboratories, One Mind, and Pfizer.

Rancho BioSciences
Posters Posters
Ontology Data Harmonization

Enhancing Biomedical Data FAIRification by Streamlining Data Harmonization with a Terminology Management Solution

Rancho BioSciences' Terminology Management Solution (TMS) combines AI-assisted semantic and fuzzy phonetic algorithms with 40+ pre-loaded biomedical ontologies to automate terminology mapping, reducing manual curation time by 50%+ and demonstrating 30% more high-quality matches than free tools when mapping real-world patient conditions to OMOP standards. In practical applications, TMS successfully mapped 10,500+ clinical terms for 20,000 de-identified patients (saving 500+ hours), achieved 65% high-quality matches for 1,875 drug names to OMOP ingredients, and enabled automated validation of AI safety predictions (saving 100+ hours), establishing itself as a critical tool for FAIRifying real-world data across drug discovery, clinical genetics, and biomedical research applications through its intuitive web interface and robust API integration capabilities.

Rancho BioSciences
Posters Posters
Single-Cell Ontology

Enhancing Single-Cell Data Integration and Discovery through Knowledge Graph Extraction from scGPT Embeddings

Rancho BioSciences developed a novel method to extract knowledge graphs from single-cell embeddings using scGPT (a large language model) applied to 11.2 million deeply curated cells from the SCDS consortium, creating nodes representing cell clusters connected by predicates ("similar to," "subset of") and enriched with disease, tissue, perturbation, and gene expression annotations. This approach successfully identified high-resolution cell subtypes across datasets—such as two distinct hepatocyte superclusters (21 and 22) differentiated by transcription factors RORA and ZBTB20 and metabolic pathways—demonstrating how knowledge graph integration bridges diseases and perturbations to genes through cell populations, enabling discovery of biologically relevant cellular heterogeneity that traditional annotations miss and accelerating insights into disease mechanisms and drug responses.

Rancho BioSciences
Posters Posters
Data Harmonization Bioinformatics

Extraction of High Content Toxicity Biomarkers Data from Scientific Literature

Rancho BioSciences and NIEHS Division of Translational Toxicology developed an automated pipeline using GPT-4, vector databases (gte-large embeddings), and trained SVC/MLP models to extract and summarize toxicity biomarker data from scientific literature, annotating >100 genes with information on transcriptional associations with disease/toxicity, upstream regulation, protein interactions, and chemical responses across multiple organs. The methodology generated single-sentence functional descriptions that, when clustered in latent embedding space, revealed tissue-spanning gene clusters involved in detoxification, immune modulation (GO:0006954 inflammatory response, FDR 0.00013), and ECM remodeling, providing a scalable AI-driven tool for toxicogenomics biomarker discovery that advances in vivo toxicity screening and characterization of chemical impacts on human health through comprehensive transcriptome-level analysis.

Rancho BioSciences
Posters Posters
Single-Cell Ontology

High-Resolution Single-Cell RNA-seq Autoimmune Atlas for Cellular Characterizationand Target Discovery in Rheumatoid Arthritis & Systemic Lupus Erythematosus

Rancho BioSciences and Nucleome Therapeutics constructed a unified single-cell RNA-seq Autoimmune Atlas integrating 1 million cells from 218 donors across 6 datasets spanning rheumatoid arthritis (RA), systemic lupus erythematosus (SLE), and controls in synovial tissue, skin, and peripheral blood, leveraging SCDS Consortium datasets and CELLxGENE Census reference data. The atlas employs multi-method cell-type annotation (scArches, CellTypist, scGPT, Azimuth), scVI batch correction with ~2,000 curated tissue-specific marker genes, and neighborhood-concordance filtering (≥55% consensus) to achieve high-resolution characterization of immune populations including dendritic cells, memory B cells, and regulatory T cells, revealing shared and tissue-specific transcriptional signatures that enable disease mechanism hypothesis generation, biomarker discovery, and therapeutic target prioritization in autoimmune pathogenesis research.

Rancho BioSciences
Posters Posters
Single-Cell Ontology

Identification and Annotation of Publicly Available Single Cell RNA-sequencing Data

Rancho BioSciences developed DataCrawler, an automated tool that identifies and annotates publicly available single-cell RNA-sequencing datasets across 12+ repositories (SRA, EGA, dbGaP, GEO, SCP, bioRxiv, PubMed) using keyword-based searches and FAIR principles, cataloging over 4,200 unique human studies with 1,800+ ex vivo clinical samples harmonized using Terminology Management Service for disease, tissue, and assay metadata. In collaboration with Cellarity, analysis revealed exponential growth in scRNA-seq/snRNA-seq publications annually, with 10x Genomics as the dominant technology, healthy/normal tissues (especially blood and bodily fluids) as the most profiled samples for reference purposes, and cancer types (including lymphomas) as the most common disease states after healthy subjects, enabling AI-driven drug discovery platforms to rapidly access and leverage public single-cell data for hypothesis testing, validation, and therapeutic target identification across all therapeutic areas.

Rancho BioSciences
Posters Posters
Multi-Omics Single-Cell

Integrating Bioinformatics and Metadata Harmonization Pipelines to Ensure High-Quality, AI-Ready Datasets in the Single-Cell Data Science Consortium Initiative

Rancho BioSciences' Single-Cell Data Science Consortium has delivered over 1,000 datasets comprising 100+ million annotated cells from 800+ publications, employing a harmonized 5-entity, 90-attribute data model aligned with public ontologies (DOID, UBERON, CL, MONDO, MeSH, EFO) and consensus cell-type annotation (popV) to create AI-ready single-cell resources across autoimmune (9M cells), cancer (24.7M cells), cardiovascular (3.5M cells), neurological (11M cells), and other disease areas. The integrated bioinformatics and metadata curation pipeline combines automated ontology mapping with expert SME review, standardized FASTQ reprocessing with batch correction, and multi-method cell annotation to generate 9 disease-specific atlases including a 1.6M-cell Digestive System Cancer atlas with ML-based malignancy scoring (CancerFinder), delivering perpetual-access R/Python-compatible datasets that accelerate reproducible science, therapeutic target discovery, and hypothesis generation across pharmaceutical research applications with SCDS2 expanding to 120+ datasets and 5 new atlases in 2025.

Rancho BioSciences
Posters Posters
Ontology Data Harmonization

Integrative Approach to Data Harmonization: Empowering Biomedical Research with Rancho Biosciences' Terminology Management Solution (TMS)

Rancho Biosciences' Terminology Management Solution (TMS) combines phonetic (Fuzzy) and AI-assisted semantic mapping algorithms to harmonize biomedical terminology across diverse datasets, demonstrating superior performance with 99% recall and 74% correct top-match accuracy when mapping disease terms to DOID—outperforming commercial tools (Tool 1: 38-47% accuracy; Tool 2: 92% recall, 69% precision). TMS enables researchers to align raw terms with public and custom ontologies (DOID, UBERON, CL, MONDO, MeSH) through an intuitive web interface or RESTful APIs, supporting 2-dimensional mapping for sample-level metadata harmonization, significantly reducing manual curation time while improving precision in term matching through cosine similarity scoring that better separates correct from incorrect mappings, positioning TMS as a critical infrastructure for biomedical data integration, interoperability, and AI-ready dataset preparation in pharmaceutical research and translational safety assessments.

Rancho BioSciences
Posters Posters
Proteomics Bioinformatics

Large scale proteomics to enable precision medicine for Alzheimer's disease

Janssen Pharmaceutica, Rancho BioSciences, and ACE Alzheimer Center Barcelona analyzed CSF proteomics from 1,321 Alzheimer's patients using SomaScan 7k assay, identifying MMP-10 as a novel prognostic biomarker that augments A/T/N framework prediction of MCI-to-dementia progression (multivariate Cox p=4.04e-05) and validating molecularly distinct Alzheimer's disease subtypes through GSVA, WGCNA, and consensus clustering methods. Analysis revealed two major pathological signatures—Neuronal Plasticity (associated with EMIF-AD MBD subtype 1, enriched for synaptic/axogenesis proteins with elevated pTau/tTau) and Blood-Brain Barrier Dysfunction (associated with subtype 3, enriched for complement/coagulation proteins with increased CSF total protein and albumin ratio, FDR p=9.20e-31)—demonstrating that large-scale proteomics enable precision patient stratification beyond A/T/N staging to optimize treatment selection, accelerate clinical trials, and support discovery of subtype-specific therapeutic targets through participation in UK Biobank Pharma Proteomics Project (62k samples, Olink 3k proteins) and Global Neurodegeneration Proteomics Consortium (40k+ samples, Somalogic 7k proteins).

Rancho BioSciences
Posters Posters
Data Harmonization Bioinformatics

Molecular-based enrichment strategy for Nectin-4 targeted Bicycle toxin conjugate BT8009

Bicycle Therapeutics, Sarah Cannon Research Institute, and Rancho BioSciences developed a molecular enrichment strategy using SDHC copy number (CN) amplification as a surrogate biomarker for Nectin-4 expression in the BT8009 Phase 1/2 clinical trial (NCT04561362), demonstrating 100% positive predictive value when SDHC CN≥3 predicts Nectin-4 IHC positivity (tumor membrane/cytoplasmic H-score ≥100) in 100 TNBC samples. Analysis of TCGA PanCancer Atlas revealed strong correlation between Nectin-4 CN and transcript expression across cancer types (Kruskal-Wallis p<0.01), with SDHC—located 225kb from Nectin-4 on chromosome 1q23 and included on FoundationOne®CDx panel—serving as a readily accessible screening tool that enables pre-identification of patients likely to have Nectin-4-positive tumors for enrollment in trials of BT8009, a Bicycle Toxin Conjugate linking Nectin-4-targeting bicyclic peptide to MMAE cytotoxin for treatment of relapsed/refractory solid tumors including bladder cancer and triple-negative breast cancer where Nectin-4 is overexpressed.

Rancho BioSciences
Posters Posters
Genomics Bioinformatics

Multiple cereblon genetic changes are associated with acquired resistance to lenalidomide or pomalidomide in multiple myeloma

University of Oxford, Bristol Myers Squibb, and Rancho BioSciences analyzed whole-genome sequencing (455 patients) and RNA-seq data (655 patients) across newly diagnosed, lenalidomide (LEN)-refractory, and pomalidomide (POM)-refractory multiple myeloma cohorts, revealing that nearly one-third (29.6%) of POM-refractory patients harbor CRBN (cereblon—the essential IMiD/CELMoD binding protein) alterations including point mutations (9.3%), copy losses/structural variations (24%), and exon-10 spliced transcripts, representing significant increases from baseline (ND: 0.5% mutations, 1.5% copy loss). All three CRBN aberration types—including previously undescribed gene copy losses and structural variants (translocations, inversions)—were independently associated with inferior progression-free survival (median PFS 2.0-2.3 vs 6.5 months, p<0.0001) and overall survival in LEN-refractory patients receiving POM-based therapy, with high exon-10 spliced transcript ratio (>2.6, deleting LEN/POM-binding region) also predicting reduced PFS in newly diagnosed patients receiving induction therapy, establishing CRBN dysregulation as the single most clinically significant contributor to IMiD resistance and informing patient selection for sequential CRBN-targeting therapies including novel CELMoDs (iberdomide, CC-92480) and PROTACs in development.

Rancho BioSciences
Posters Posters
Bioinformatics Clinical Data

NCATS FRDB: Fast Response Database for Drug Repositioning

Rancho BioSciences and NCATS developed the Fast Response Database (FRDB), a publicly accessible web resource (https://drugs.ncats.io) that augments the NCATS Inxight Drugs portal with manually curated pharmacokinetic data (Cmax, AUC, t1/2, fraction unbound for 567 drugs), adverse events/toxicity profiles (53,600 records including dose-limiting toxicities, MTDs, discontinuation events for 2,560 drugs), drug-drug interactions (15,500 records with clinical evidence for 1,400 drugs), and sourcing information to accelerate emergency drug repositioning and PK/PD modeling. The database enables dose-concentration relationship modeling with 22% mean relative error in AUC extrapolation, anti-target identification (e.g., hERG, 5-HT2C, VEGFR2 selectively inhibited at toxic vs safe doses), and rapid assessment of whether investigational or approved drugs can achieve required target concentrations safely in patients—critical capabilities for responding to healthcare emergencies by repurposing existing therapeutics with known safety profiles across multiple administration routes (oral n=1,017; IV n=139; topical, respiratory, subcutaneous) and clinical populations.

Rancho BioSciences
Posters Posters
Data Harmonization Bioinformatics

NCATS Stitcher: Data Integration and Guided Curation Tool

Rancho BioSciences and NCATS developed Stitcher, a RESTful API (stitcher.ncats.io/api) and curation web interface that integrates 13 heterogeneous drug data sources (ClinicalTrials.gov: 305,805 entries; G-SRS: 100,264; DailyMed: 167,689; DrugBank: 11,922; plus FDA, Broad Institute, and manually curated datasets) into a Neo4j knowledge graph using a novel deterministic algorithm that guarantees accuracy of entity connections while automatically detecting data inconsistencies like shared UNII clashes, orphan substances, and erroneous linkages. The system addresses critical drug repurposing challenges—where research has increased 35-fold from 2010 to 2018—by combining incomplete data sources to identify ground truth, enabling non-destructive manual curation guided by automated quality tests, maintaining complete data provenance, and resolving entity normalization issues that plague traditional ETL tools, ultimately making previously untapped Big Data (98% of potentially useful data remains unanalyzed) accessible for translational research and therapeutic discovery through a comprehensive substance knowledge graph that feeds corrections back to original data sources.

Rancho BioSciences
Posters Posters
Data Harmonization Bioinformatics

NF Research Tools Central: A disease-specific knowledgebase of experimental tools

Sage Bionetworks, Rancho Biosciences, and Gilbert Family Foundation developed NF Research Tools Central (https://tools.nf.synapse.org), an open-access database and web portal cataloging 1,000+ neurofibromatosis type 1 (NF1) and RAS-relevant research tools including animal models, cell lines, antibodies, genetic reagents, and biobanks, integrating data from Cellosaurus, AntibodyRegistry, RRID Portal, literature, and community contributions. The platform features a relational MySQL database built using the schematic Python library and deployed on Synapse.org with a React-based interface that enables filtering by standardized metadata, provides detailed "Tool Detail" pages with vendor links and experimental observations (pathology, usage notes), and incorporates AI-assisted curation using large language models to generate comprehensive tool descriptions—addressing critical gaps in existing generalist databases by combining disease-specific focus, in-development models, observational data, and streamlined community submission processes to help NF1 researchers identify, obtain, and effectively use appropriate experimental tools supported by linked datasets generated from those resources.

Rancho BioSciences
Posters Posters
Genomics Spatial Transcriptomics

OncoDM and Spatial Innovation Initiatives Provide High-Quality, FAIR and AI-Ready Data to Fast-Track Drug Discovery Insights

Rancho BioSciences developed two precompetitive initiatives delivering F.A.I.R. (Findable, Accessible, Interoperable, Reusable) and AI-ready datasets: (1) Oncology Data Mart (OncoDM) providing uniformly processed bulk RNA-seq data from 38,804 samples across 133+ datasets (TCGA: 11,400 cancer samples; GTEx: 18,000 normal reference; GEO: 9,404 across various cancers) with differential gene expression analysis and custom outlier detection, and (2) Spatial Innovation Initiative (SII) delivering 43 high-quality spatial transcriptomics datasets (10x Visium, Nanostring GeoMx/CosMx, MERSCOPE platforms) from human and mouse tissues across healthy, cancer, and disease states with extended spatial-specific data models. Both initiatives feature harmonized metadata aligned to custom data models, consistent bioinformatics pipelines starting from raw FASTQ files with standardized genome mapping/quantification, and integration-ready formats enabling seamless cross-dataset comparison, computational pipeline integration, and AI/ML model training—with similar initiatives expanding to neurology, immunology, and gastrointestinal disease areas to accelerate drug discovery through cost-shared, collaborative value delivery that unlocks insights from disparate public domain data.

Rancho BioSciences
Posters Posters
Single-Cell AI

Quality Data and Creative Approaches Drive GenAI's Transformative Journey in Biopharma

Rancho Biosciences demonstrates how high-quality, harmonized AI-ready data drives breakthrough GenAI applications including multi-agent pipelines reducing toxicology literature summarization from >16 to 1.5 hours per gene, scGPT foundational models enabling rapid rare cell-type discovery across 11.2 million SCDS cells spanning cancer/neurological/immune diseases, and validation that proper data curation is essential to unlock McKinsey's estimated $60-110 billion annual pharma value from AI in R&D, clinical operations, and IT—with pharma's 200M+ cell datasets (~8 trillion tokens, exceeding GPT-3 training) representing untapped opportunities for foundational model development despite lower compute investment versus big tech.

Rancho BioSciences
Posters Posters
Proteomics Bioinformatics

Quantitation of CD137 and Nectin-4 expression across multiple tumor types to support indication selection for BT7480, a Bicycle tumor-targeted immune cell agonist (Bicycle TICA)

Bicycle Therapeutics and Rancho BioSciences established a translational pipeline combining TCGA RNA-seq analysis (~10,000 samples, 36 cancer types) with 19-plex MultiOmyx™ immunofluorescence spatial proteomics of 43 FFPE tumor samples (HNSCC, lung, bladder, breast) to quantify co-expression of CD137 (immune agonist target) and Nectin-4 (tumor antigen) for BT7480, a Bicycle TICA™ linking both targets. Analysis revealed 50-78% of samples co-expressed CD137 and Nectin-4 at clinically relevant levels (>1% positive cells), with CD137+ immune infiltrates predominantly comprising CD4+ T cells (32-47%), CD8+ T cells (16-23%), and macrophages (2-8%), and spatial analysis demonstrating CD137+ cells localize within 150 microns of Nectin-4+ tumor cells across indications, supporting prioritization of head & neck, lung adenocarcinoma/squamous, and bladder cancers for BT7480 clinical development and validating MultiOmyx™ as a companion diagnostic for proof-of-mechanism monitoring in first-in-human trials.

Rancho BioSciences
Posters Posters
Genomics Data Harmonization

Rare Disease-Target Gene Treatment Knowledge Mining Partnership between Rancho BioSciences and Rady Children's Institute for Genomic Medicine (RCIGM)

Rancho BioSciences and Rady Children's Institute for Genomic Medicine developed a Genome-to-Treatment platform (GTRx/BeginNGS) through systematic knowledge mining of public data sources, curating 1,040 disease-gene pairs and identifying 18,003 potential interventions from PubMed, clinical trials, and conference abstracts, with clinical genetics teams retaining 2,043 (11.3%) evidence-based treatments for ICU and newborn screening applications after Delphi panel review for efficacy, pediatric appropriateness, and quality of evidence. The partnership reduced time per gene-disease pair from >15 hours (2020) to <5 hours through automation, enabling rapid access to curated treatment recommendations at physicians' fingertips—directly benefiting all 28 (3.7%) of 750 healthy newborns identified with rare diseases (including phenylketonuria, cystic fibrosis, mucopolysaccharidoses, metabolic disorders) in an ongoing Phase 3 clinical trial, reducing time from diagnosis to intervention from hours/days to minutes through a harmonized REDCap-based database deployed across 20+ medical centers including UCSD, Scripps, Kaiser Permanente, and U.S. Naval hospitals.

Rancho BioSciences
Posters Posters
Single-Cell Bioinformatics

Single-Cell RNA-SEQ of hPSC-derived neural cell types reveals early glia-specific differentiation trajectories

NCATS and Rancho Biosciences performed single-cell RNA-seq on 12,771 cells including hPSC-derived astrocytes (generated via direct RGC-to-astrocyte conversion bypassing neurogenesis), commercial astrocytes, and cortical glutamatergic neurons, identifying 11 transcriptionally distinct clusters and three slingshot pseudotime trajectories mapping RGC differentiation into two astrocyte branches and one neuronal branch. Integration with human fetal cortex data (Nowakowski 2017, gestational weeks 5-37) and trajectory-based differential expression analysis (tradeSeq) revealed novel astrocyte-specific markers (SOD1, CLU, ATP6AP2, CD44, NFIA) and gene expression modules characterizing direct RGC-to-astroglia transition, providing insights into the "gliogenic switch," cellular diversity, and astrocyte maturation mechanisms relevant to neurodegenerative disease research through chemically-defined differentiation that achieves robust NFIA/CD133 upregulation and GFAP+ stellate astrocyte morphologies while eliminating neuronal marker TUBB3.

Rancho BioSciences
Posters Posters
Single-Cell Ontology

Single-Cell Data Science Consortium Enables Rapid Analysis of High Value Public Datasets

Rancho BioSciences leads the Single Cell Data Science (SCDS) Consortium, a pre-competitive initiative that grew from 4 charter members (2022) to 7 members (2023), delivering 249 analysis-ready datasets comprising 32 million harmonized cells (14M+ healthy, 4.7M cancer, 1.5M immune/autoimmune, 1.4M neurodegenerative) with metadata curated to a 6-entity, 112-attribute data model mapped to official ontologies and provided in multiple formats (Seurat RDS, scanpy h5ad, CSV). The consortium addresses single-cell data challenges—lack of standardization, batch effects, algorithm explosion, integration complexity—through automated cell-type annotation (CellTypist, scArches achieving 24-77% coverage supplementing 48.5% author-provided labels), optimized integration methods combating batch effects, and construction of disease/tissue-specific atlases (autoimmune, cancer, neurological) from top tissues (brain, lung, blood, bone marrow, kidney) across 3,522 donors and 14,837 samples, accelerating reproducible science and drug discovery insights by making high-dimensional public domain data analysis-ready for computational pipelines and cross-dataset comparison.

Rancho BioSciences
Posters Posters
Single-Cell Ontology

Single-Cell Data Science Consortium Enables Rapid Analysis of High Value Public Datasets

Rancho BioSciences leads the Single Cell Data Science (SCDS) Consortium, a pre-competitive initiative launched in 2022 with 4 charter members (growing to 6 by 2023), delivering 115 analysis-ready datasets from 96 studies comprising 17.8 million cells across 6,885 donors with metadata harmonized to a 4-entity, 75-attribute data model mapped to official ontologies and provided in three formats (Seurat RDS, scanpy anndata, CSV). The consortium ingested 7.5 million diseased cells including cancer (49%: lung, hematological, GI), neurological diseases (19%: HD, PD, AD), autoimmune conditions (15%: psoriatic arthritis, ulcerative colitis, dermatitis), and GI dysfunction (45%), addressing pharma challenges of lack of standardization, batch effects, and integration complexity through a dataset tracker enabling member prioritization, with Year 2 plans including migration to Python-centric pipeline for scalability, comprehensive automated cell-type annotation using reference/ML tools mapped to Cell Ontology (supplementing 49.1% author-provided annotations), creation of tissue/disease-specific atlases, and expanded versioning/logistics support to accelerate drug discovery through reproducible analysis of high-dimensional public domain single-cell data.

Rancho BioSciences
Posters Posters
Single-Cell Ontology

Single-Cell Data Science Consortium Enables Rapid Analysis of High Value Public Datasets

Rancho BioSciences leads the Single Cell Data Science (SCDS) Consortium, a pre-competitive initiative that grew from 4 charter members (2022) to 6 members (2023), delivering 168 analysis-ready datasets from 159 studies comprising 25 million harmonized cells (11.3M healthy, 4.7M cancer, 1.5M immune/autoimmune, 1.4M neurodegenerative) across 11,529 donors with metadata curated to a 4-entity, 99-attribute data model mapped to official ontologies and provided in three formats (Seurat RDS, scanpy h5ad, CSV). Year 2 updates include automated cell-type annotation using CellTypist and scArches (supplementing author-provided labels that cover 24-76% of cells), construction of disease-specific atlases (autoimmune atlas integrating systemic scleroderma datasets with optimized batch-effect correction), and coverage of >500k cells each from blood, lung, liver, heart, colon tissues and >250k cells from skin, bone marrow, lymph nodes, addressing pharma challenges of standardization, batch effects, and integration complexity to accelerate drug discovery through reproducible analysis of high-dimensional public domain single-cell data.

Rancho BioSciences
Posters Posters
Spatial Transcriptomics Ontology

Spatial Innovation Initiative Delivering Harmonized Analysis-Ready Datasets

Rancho BioSciences launched the Spatial Innovation Initiative in 2024 with three charter members to make spatial transcriptomic data accessible through harmonized, analysis-ready datasets across five technology platforms: 10x Visium (50μm spots, 16k genes, 40mm²), Nanostring GeoMx/CosMx (multi-cell to single-cell, 11k-1000 genes), Vizgen MERSCOPE (single-cell imaging, 300-500 genes, 100mm²), and 10x Xenium (subcellular imaging, 200-400 genes, 1400mm²). Using DataCrawler to identify public datasets (predominantly from GEO repository, with 10x Visium as most popular platform showing exponential publication growth), Rancho curates metadata to an expanded 76-attribute, 5-entity data model (building on transcriptomic models with spatial-specific fields) mapped to public ontologies, delivering Seurat RDS objects, Python h5ad anndata files, spatially variable gene tables, QC plots, and metadata spreadsheets that enable analysis of tissue organization, cellular interactions, and spatial expression patterns (e.g., polarized vs. random marker distribution in healthy vs. diseased gut, localized UMI expression in breast tissue clusters), with 10 priority datasets per member being processed following successful pilot validation across healthy/diseased tissues in multiple formats.

Rancho BioSciences
Posters Posters
Spatial Transcriptomics Ontology

Spatial Innovation Initiative Delivering Harmonized Analysis-Ready Datasets

Rancho BioSciences' Spatial Innovation Initiative (launched 2024, 3 charter members) delivers analysis-ready spatial transcriptomic datasets across five platforms (10x Visium, Nanostring GeoMx/CosMx, Vizgen MERSCOPE, 10x Xenium) with metadata harmonized to a 76-attribute spatial-specific data model mapped to public ontologies, providing Seurat RDS, Python h5ad, spatially variable gene tables, and QC plots—validated through pilot analysis of healthy/diseased tissues (breast, stomach, gut) showing polarized vs. random expression patterns—with 10 priority datasets per member being processed from rapidly growing public repositories (predominantly GEO) to enable insights into tissue organization, cellular interactions, and spatial gene expression for drug discovery.

Rancho BioSciences
Posters Posters
Data Harmonization Bioinformatics

Streamlined Flow Cytometry Data QC Through Interactive Visualization

Rancho BioSciences, Genentech Computational Catalysts, and Translational Medicine OMNI developed FlowinSight, an automated R-based interactive visualization application that integrates clinical data (SDTM, ADaM), high-dimensional flow cytometry biomarker data from CROs, and assay metadata to enable real-time quality control of multi-color flow cytometry assays in early-phase clinical trials. The platform performs automated QC checks (file/sample reconciliation, data integration validation, outlier detection) identifying issues like missing raw data, duplicate samples, gating errors, visit/date mismatches, and missing reportables, presenting results through interactive plots with hover metadata, sample tables with QC status tracking via Google Sheets API, comment forms for assigning QC status, and integrated raw data PDF viewer for gating information—enabling scientists and operations leads to rapidly identify data quality issues, detect unexpected longitudinal trends from baseline, and take swift corrective actions through study-agnostic R packages with CRON-scheduled updates that streamline assessment of therapy impact on immune cell populations for dosage efficacy and safety evaluation.

Rancho BioSciences
Posters Posters
Ontology Data Harmonization

Towards a comprehensive view of diagnoses in UK Biobank by data curation and aggregation

Rancho BioSciences and Takeda developed comprehensive curation workflows to integrate UK Biobank GP (primary care) data (230,105 participants, 123.7M clinical records; 222,122 participants, 57.7M prescription records) with hospital inpatient (HESIN), cancer registry, and self-reported diagnoses, addressing challenges of incomplete READ-to-ICD10/OPCS4 mappings, one-to-many code relationships, and distorted formats through combined automated TRUD mapping and manual curation.

Rancho BioSciences
Posters Posters
Multi-Omics FAIR Data

Unleashing Data's Full Potential through Integrated Data Marts

Rancho BioSciences and Genentech/Roche developed standardized end-to-end workflows integrating clinical and high-dimensional biomarker data (genomics, transcriptomics, proteomics, imaging) into MultiAssayExperiment (MAE) objects within 20 deployed data marts across therapeutic indications, enabling 54 key R&D insights since 2020 through FAIR-ification (Findable, Accessible, Interoperable, Reusable) principles.

Rancho BioSciences
Posters Posters
Ontology AI

Using Artificial Intelligence to Enhance Terminology Mapping Workflow from Data Collection to Standardization

Rancho BioSciences evaluated AI-assisted semantic and phonetic mapping algorithms for terminology standardization by comparing automated ICD10CM mapping against manual curation of 22,000 UK Biobank READ code descriptions, testing four embedding approaches: Rancho's phonetic Fuzzy tool (PostgreSQL pg_trgm, 33% top-1 accuracy), OpenAI ada-2 (1536-dimension embeddings, 33%/62.6%/93.4% top-1/5/100 accuracy), GTE-large contrastive learning model (1024-dimension, 36.4%/69.6%/90.2%), and Instructor-xl instruction-finetuned model (768-dimension, 30.1%/55.2%/85.3%). Three hierarchical algorithms were tested to leverage ontology structure: greedy top-down (13.9% accuracy, prone to wrong-branch errors), N-max hierarchical top-down (19.3%, non-monotonous similarity issues), and greedy bottom-up (37.4%, best performance but imprecise semantic mapping)—demonstrating that while embedding quality significantly impacts results and hierarchical approaches provide modest improvement over reference nearest-neighbor methods, complex semi-automated solutions combining AI-assisted mapping with manual expert curation remain necessary for FAIR data normalization, error correction, and alignment to standard ontologies (ICD, SNOMED CT) enabling findable, integrable, and reusable clinical/EHR data across diverse sources and formats.

Rancho BioSciences
Posters Posters
Ontology AI

Using Graph Technologies and Artificial Intelligence to Power Terminology Management

Rancho BioSciences developed an integrated terminology management ecosystem combining three tools: (1) Data Crawler—a web application crawling PubMed, ClinicalTrials.gov, GEO, SRA, EGA, ArrayExpress, DbGaP (with SCP, FigShare, Zenodo forthcoming) that extracts study/sample-level metadata, annotates free-text fields against standard ontologies (Uberon, DOID, BTO), calculates journal ranks/citation indices, and enables incremental updates; (2) Terminology Mapping—using PostgreSQL trigram indexing for phonetic "fuzzy" mapping (50-75% successful mapping) combined with OpenAI embeddings and SciGraph (Neo4j-based ontology store) for semantic alignment, common ancestor searches, and hierarchical queries across public ontologies (OMOP, SNOMED, DOID, NCIT, MedDRA); and (3) Categorization Tool—leveraging OpenAI embeddings with DBSCAN clustering (density-based, no predetermined cluster count, user-adjustable epsilon distance) to QC data dictionaries, test heterogeneity, suggest term groups/domains, and identify outliers for optimal data organization. The all-in-one solution enables FAIR data principles (findable, integrable, reusable) through automated normalization, error correction, and alignment to well-established ontologies, significantly reducing manual curation effort for harmonizing large data streams from diverse public resources.

Rancho BioSciences
Posters Posters
Ontology AI

Using Literature-Based Knowledge Extraction to Develop a Disease Ontology for Skin Dysbiosis

Rancho BioSciences developed a custom Skin Dysbiosis ontology by using DataCrawler to identify ~500 relevant papers, extracting subject-predicate-object tuples through three NLP pipelines (SciSpacy-FT, SciSpacy-abstract, PubTator), and mapping extracted entities (genes, proteins, chemicals, phenotypes) to public ontologies (UniProt, NCBI, BTO, PubChem, MeSH, NCIT, GO, CL, DOID, CHEBI) using TMS with ~70% accuracy requiring minimal manual revision. The workflow generated Biolink predicates using GPT-4 (86% precision, 20% recall for relationship extraction), built the ontology in R with n-triples/turtle/OWL formats via rapper and Protégé, and created an interactive knowledge graph connecting 2,244+ tuples (e.g., "Th17 cells associated_with psoriasis," "M. restricta impairs skin barrier") across domains (biological processes, cell types, chemicals, diseases, genes, tissues, fungal structures, species)—demonstrating proof-of-concept for DataCrawler and TMS as effective tools for literature-based knowledge extraction that advances understanding of skin dysbiosis pathophysiology, facilitates computational reasoning through hierarchical ontology relationships, enables complex hypothesis-testing queries, and informs improved treatment strategies through systematic organization of disparate biomedical terms into FAIR, machine-readable formats.

Rancho BioSciences
Posters Posters
Ontology Data Harmonization

Utilizing NIIMBL and ISA-88 Framework to Create a Proof of Concept Pharmaceutical Manufacturing Ontology

Rancho BioSciences and Merck developed a proof-of-concept pharmaceutical manufacturing ontology by leveraging the NIIMBL (National Institute for Innovation in Manufacturing Biopharmaceuticals) ontology framework—built on BFO, IOF, and QUDT backbones with ISA-88 manufacturing process standards—to model Merck's manufacturing processes across different products, scales, and production sites, using tablet compression as the pilot modality. The project produced Conceptual and Logical Data Models (CDM/LDM) defining entity information, sample attributes, and valid predicates based on standardized NIIMBL and ISA-88 terminologies (parameter, process, procedure hierarchies), creating a reusable data structure that visualizes pharmaceutical manufacturing processes dynamically, demonstrates how production processes are monitored and controlled, facilitates tech transfer during scale-up, and enables application across multiple manufacturing modalities through improved data interoperability, semantic integration, inference/reasoning capabilities, and data integrity validation by establishing common language and hierarchical relationships between key manufacturing steps, processes, equipment, and parameters in a knowledge graph format rather than traditional documents or databases.

Rancho BioSciences
Posters Posters
Bioinformatics Data Curation

Streamlined Flow Cytometry Data QC Through Interactive Visualization

Rancho BioSciences, Genentech Computational Catalysts, and Translational Medicine OMNI developed FlowinSight, an automated R-based interactive visualization platform that streamlines flow cytometry quality control in early-phase clinical trials by integrating clinical data (SDTM, ADaM), high-dimensional biomarker data from CROs, and assay metadata through three-stage automated QC checks (file/sample reconciliation, integration validation, visual data QC) identifying data mismatches, missing files, gating errors, and longitudinal trends. The application features interactive plots with hover metadata, QC summary dashboards, Google Sheets API-integrated sample tracking, comment forms for QC status assignment, and integrated raw data PDF viewer for instant gating information access—enabling scientists and operations leads to rapidly detect outliers, assess dosage efficacy and safety, and take corrective actions in real-time through 4 study-agnostic R packages with CRON-scheduled updates that reduce manual review time and accelerate clinical trial decision-making for multi-color flow cytometry assessment of therapy effects on immune cell populations.

Rancho BioSciences
Posters Posters
Multi-Omics Single-Cell

Data Alliances: A collaborative opportunity for FAIR and AI-ready data

Rancho BioSciences leads four collaborative Data Alliance initiatives providing harmonized, AI-ready biomedical datasets: (1) Oncology Data Mart integrates bulk RNA-seq from TCGA (33 cancer types), GTEx, and GEO with standardized processing; (2) Single Cell Data Science Consortium delivers 112 million reprocessed cells across 1,154 datasets (h5ad, Seurat, TileDB) with 100% cell-type annotation and 14 disease atlases; (3) Spatial Innovation Initiative provides 33 spatial transcriptomics datasets across platforms (10x Visium, GeoMx, CosMx, MerFish) spanning oncology, immunology, and neurodegeneration; (4) Perturb-Seq Initiative targets 180+ studies with genetic (CRISPR) and chemical perturbations enabling causal discovery and AI model training—all delivered as FAIR-compliant, platform-agnostic datasets through shared investment models maximizing ROI with perpetual access and no maintenance costs.

Rancho BioSciences
Posters Posters
AI Multi-Omics

Curating the Huntingtin Interactome with Claude Code as an Agentic Extraction Workflow

Rancho BioSciences, Bridlewood Consulting, and CHDI Foundation developed an agentic AI workflow using Claude Code (Claude 4.6 Opus) with iterative skill-based extraction to curate the Huntingtin Interactome (HINT) within HDinHD—an open HD research portal providing gene-centric access to omics studies, perturbation data, animal models, gene set enrichment libraries, and literature-derived HTT protein-protein interactions across 70 structured fields per experiment. After testing Google NotebookLM (variable format, hallucinations), OpenAI GPT-5 (frequent timeouts, >15 min/article), Claude Code was selected for its consistency and speed (~2 min/article dual-run extraction) using a five-phase pipeline (Discovery → Survey → Commitment → Extraction → Validation) with three mandatory QC gates and structured audit logs, refined through 29 iterative cycles incorporating SME feedback via /skill refine, /skill compare-results, and /skill sme-compare commands.

Rancho BioSciences
Posters Posters
Data Harmonization AI-Ready Data

Unlocking Perturb-Seq through Harmonization and Scalable Exploration

Rancho BioSciences addresses drug discovery challenges in target identification by curating, harmonizing, and processing Perturb-seq (pooled genetic/chemical perturbations with single-cell RNA-seq profiling) and spatial transcriptomics datasets (gene expression with preserved tissue architecture) through a three-stage workflow: (1) Find Relevant Data using DataCrawler to search 12 public repositories, (2) Standardize Metadata with Terminology Management Solution mapping to desired ontologies and data models, and (3) Bioinformatics Processing uniformly processing molecular data to enable cross-study comparison—delivering AI- and ML-ready datasets that overcome inconsistent formats, incomplete metadata, and lack of standardization barriers.

Rancho BioSciences
Posters Posters
Clinical Data Clinical Development

FlowInsight: Automation of Clinical Flow Cytometry Data Workflows for Enhanced Delivery and Traceability

Rancho BioSciences and Genentech developed FlowInsight, a two-part solution integrating fragmented flow cytometry data from Contract Research Organizations (CROs) through: (1) ETL Pipeline that searches, standardizes, reconciles sample assays across sources (fcs files, PDFs, tabular datasets), performs automatic QC checks, and outputs analysis-ready datasets via six-stage process (extract study metadata/raw data → map to standard format → combine unique sample assays prioritizing highest quality source → integrate with clinical metadata [patient cohort, site] → validate fields → output to database), and (2) FlowInsight RShiny Application with three modules enabling scientific QC, analysis, and operational tracking.

Rancho BioSciences
Blog Blog
Blog
AI and Analysis-ready Data Data Curation and Management AI Solutions and Augmentation Bioinformatics & Data Science

Same Problem, Different Teams. What I Heard at BioIT 2026

I spent three days at BioIT World in Boston a few weeks ago. As I spoke with data scientists, computational biologists, ...

Candace Ruff, Sr. Product Manager
Blog Blog
Blog
AI and Analysis-ready Data Data Curation and Management Data Quality Bioinformatics & Data Science

The 'O' word

Not so long ago, there was a period where we quietly stopped using the word "ontology" in sales conversations for fear ...

Jane Lomax