Single-cell transcriptomics · Agentic AI

scACORN

single-cell Agentic Context ORchestration Network

Context-engineered agent orchestration of specialized small language models for single-cell transcriptomic interpretation

Arash Rasti-Meymandi, Sepideh Nahali, Eustache Paramithiotis, Angela M. Cheung & Elham Dolatabadi

01 · The approach

Build local expertise.
Compose it in context.

Single-cell atlases now exceed 66 million cells, but a ranked expression profile is not yet a reliable biological answer.

scACORN separates the acquisition of domain expertise from the policy used to select and integrate it. Compact experts learn tissue-specific transcriptomic structure and grounded biological completions. A fixed language-model agent then selects and combines those experts using a natural-language playbook optimized from textual feedback.

Figure 1 scACORN constructs specialized scSLMs, learns how to route among them, and synthesizes profile-linked molecular evidence.
  1. 01

    Domain alignment

    Fit the backbone to transcriptomic geometry

    Contrastive LoRA adaptation organizes the cell-to-text backbone around a target tissue while keeping the base model frozen.

  2. 02

    Expert specialization

    Learn answers without erasing structure

    Question-conditioned completions are trained with pointwise, distributional, and relational preservation objectives.

  3. 03

    Textual orchestration

    Route, combine, and ground

    A fixed orchestrator selects one or more experts and improves its inference policy through a textual playbook, without gradient updates.

02 · Key results

Specialization changes the signal.

Across ten Tabula Sapiens tissues, each stage improves the part of the system it is designed to control: representation, expert reasoning, and evidence use.

Stage 1

Domain alignment recovers cell-type structure

Adaptation moves the language-model backbone into the range of purpose-built ranked-gene representations while preserving a reusable model state.

Macro-F10.36 0.64
Recall@50.87 0.97
Figure 2 Classification, retrieval, and UMAP structure after domain alignment.
Stage 2

Compact experts improve labels and molecular grounding

Tissue-specialized scSLMs outperform much larger general-purpose and transcriptomic models, particularly on evidence and exclusionary markers.

Mean exact match0.89
Evidence-gene F10.82
Negative-marker F10.95
Figure 3 Ontology credit, evidence-gene F1, macro-F1, and negative-marker F1.
Stage 3

A textual playbook makes evidence use more reliable

The largest optimization effect is on whether cited genes are actually supported by the input. The pattern holds under both GPT-5.4-mini and Claude Sonnet 4.5 orchestrators.

14.5%initial
3.5%optimized

Unsupported support-gene citation rate, GPT-5.4-mini

Figure 4 Playbook optimization trajectories under two fixed orchestrators.

03 · Table 1

Textual orchestration performance.

Performance across orchestrators, question types, and playbook conditions. Values are means across held-out questions; higher is better.

80

held-out questions

38
single-expert
42
compositional
CategoryModelOntology Evidence align. Marker cov. Nuanced
a Single-expert questions n = 38
Open-sourceMistral-Small-3.2-24B-Instruct0.5930.3110.5060.435
Qwen3-14B0.4600.2000.3420.371
Llama-3.1-8B-Instruct0.2160.1350.2960.273
Closed-sourceGPT-4o-mini0.4800.3080.3880.400
GPT-50.6330.2900.8650.486
SpecializedCell-o10.4110.2590.5560.476
scACORN GPT-5.4-mini; empty0.7800.5370.9020.699
scACORN GPT-5.4-mini; optimized0.8320.6320.9040.747
scACORN Claude Sonnet 4.5; empty0.7510.5430.9140.709
scACORN Claude Sonnet 4.5; optimized0.8410.6780.9220.753
b Compositional questions n = 42
Open-sourceMistral-Small-3.2-24B-Instruct0.4910.3010.4610.377
Qwen3-14B0.5690.1930.3750.381
Llama-3.1-8B-Instruct0.3380.1130.2090.280
Closed-sourceGPT-4o-mini0.5430.2570.3340.384
GPT-50.7730.3640.8070.520
SpecializedCell-o10.4990.3050.5440.482
scACORN GPT-5.4-mini; empty0.6140.0190.7930.513
scACORN GPT-5.4-mini; optimized0.8530.5080.7020.741
scACORN Claude Sonnet 4.5; empty0.6220.1390.8080.553
scACORN Claude Sonnet 4.5; optimized0.8580.5350.7160.744
Table 1 Textual orchestration performance across orchestrators, question types and playbook conditions. Evidence alignment is weighted set-F1 agreement; marker coverage is weighted recall of reference supporting and negative markers.

04 · Qualitative interactions

From ranked genes to biological reasoning.

Held-out examples from the paper appendix. Ranked-gene lists are omitted here, as in the manuscript, to focus on the question, scACORN response, and reference answer.

High agreement · Single cellEndothelial versus mesenchymal identity
User

I pasted a ranked gene list above and I am not sure what cell type this profile represents. Could you suggest the most likely cell identity and point to which markers in the list support that assignment? I am trying to decide between an endothelial versus a mesenchymal or smooth-muscle interpretation.

scACORN

Most likely identity: cardiac endothelial cell (predicted tissue: heart). This assignment is supported by endothelial markers present in the ranked list, including VWF, ENG, and EMCN. Mesenchymal and smooth-muscle alternatives are weakened by the absence of COL1A1, DCN, ACTA2, and TAGLN from the ranked profile. Confidence: high. Plausible alternative labels include generic endothelial cell, vascular endothelial cell, capillary endothelial cell, and vein endothelial cell.

Reference answer

Most likely cell identity: cardiac endothelial cell. This assignment is supported by VWF and EMCN being present in the ranked genes, with ENG further reinforcing an endothelial signature. The lack of COL1A1 provides little support for a fibroblast interpretation.

High agreement · Multi-cellImmune versus myeloid profiles
User

I pasted two ranked gene lists above corresponding to profile IDs A and B. For each profile ID, provide a single best cell-type annotation with a confidence level and short rationale, and explicitly list the supporting evidence and negative markers used from each list.

scACORN

Profile A: naive thymus-derived CD4-positive, alpha-beta T cell (confidence: high). The ranked list contains IL7R and LTB, which support a naive CD4 T-cell interpretation, while effector and activation alternatives are weakened by the absence of GZMB, IFNG, IL4, and IL5.

Profile B: mononuclear phagocyte (confidence: moderate). Antigen presentation markers including CD74, HLA-DRA, HLA-DRB1, and CST3 support an antigen-presenting mononuclear phagocyte state. B-cell and activated T-cell alternatives are weakened by the absence of canonical lymphoid and cytotoxic markers.

Reference answer

Profile A: naive thymus-derived CD4-positive, alpha-beta T cell, high confidence. IL7R and LTB are present, while canonical effector or activation markers including IFNG, IL4, IL5, and GZMB are not detected among the top-ranked genes.

Profile B: mononuclear phagocyte, high confidence. AIF1, TYROBP, and FCER1G appear in the ranked list, while lymphoid and endothelial markers including CD3D, MS4A1, PECAM1, and EPCAM are absent or low.

Biologically distinctive · Single cellRetinal cell identity
User

I pasted a ranked gene list from a cluster above. Can you give the most likely cell identity and one or two plausible alternative annotations, with a brief explanation of the evidence?

scACORN

Most likely identity: eye photoreceptor cell (predicted tissue: eye). This assignment is supported by neuronal genes SNAP25 and MAP2 in the ranked list. S100B, TMEM119, and COL1A1 are absent, weakening glial and fibroblast interpretations. Plausible alternatives include a generic retinal neuron or retinal bipolar neuron. Confidence: moderate because the detected neuronal genes are not fully photoreceptor specific.

Reference answer

Most likely identity: eye photoreceptor cell. MAP2 and SNAP25 support a neuronal retinal identity. The absence of glial or stromal markers such as S100B and COL1A1 argues against a glial or fibroblast interpretation. Plausible alternatives include other retinal neurons such as bipolar or amacrine cells because MAP2 and SNAP25 are neuronal but not fully photoreceptor specific.

Biologically distinctive · Multi-cellEpithelial versus perivascular profiles
User

I pasted two ranked gene lists above for profiles A and B. Could you annotate each profile and determine whether they represent related states within one lineage or clearly distinct cell identities?

scACORN

Profile A: enterocyte of the epithelium proper of the ileum. The ranked genes include APOA1, APOA4, SLC5A1, FABP2, and ANPEP, supporting an absorptive ileal enterocyte identity. Confidence: high.

Profile B: pericyte. The ranked genes include RGS5, NOTCH3, SPARC, IGFBP7, and CALD1, supporting a perivascular mesenchymal identity. Confidence: moderate.

The two profiles represent clearly distinct cell identities rather than related states within one lineage.

Reference answer

Profile A: enterocyte of the epithelium proper of the ileum. APOA1, APOA4, and SLC5A1 support an absorptive enterocyte identity, whereas MUC2 and CHGA are not enriched and argue against goblet or enteroendocrine interpretations.

Profile B: pericyte. RGS5, NOTCH3, and COL4A1 support a pericyte or perivascular program, whereas PECAM1 and VWF are not enriched and argue against an endothelial identity.

The two profiles are clearly distinct cell identities rather than related states of one lineage.

05 · Abstract

Single-cell atlases now exceed 66 million cells, but turning a ranked expression profile and a free-form biological question into a reliable, evidence-grounded answer remains unsolved.

scACORN combines specialized small language models with context-engineered agent orchestration for their selection and composition at inference time. Domain-aligned contrastive adaptation fits a pretrained cell-to-text backbone to target transcriptomic geometry; geometry-preserving specialization learns biological completions without eroding that geometry; and a fixed orchestrator selects and combines experts under a playbook optimized from textual feedback.

The results support specialization and orchestration as complementary responses to the heterogeneity and evidentiary demands of single-cell analysis.

06 · Paper & code

Reproduce the workflow.

The public repository includes domain alignment, expert specialization, textual orchestration, benchmark evaluation, dataset preparation, and manuscript figure generation.

Cite scACORN

@article{rasti-meymandi_scacorn,
  title = {scACORN: Context-engineered agent orchestration of specialized small language models for single-cell transcriptomic interpretation},
  author = {Rasti-Meymandi, Arash and Nahali, Sepideh and Paramithiotis, Eustache and Cheung, Angela M. and Dolatabadi, Elham},
  note = {Preprint}
}

Publication details will be updated with the preprint record.

Authors & affiliations

Arash Rasti-Meymandi1,5 · Sepideh Nahali2 · Eustache Paramithiotis4 · Angela M. Cheung1 · Elham Dolatabadi3,5

  1. 1 Department of Medicine, University of Toronto
  2. 2 School of Health Policy and Management, Faculty of Health, York University
  3. 3 York University
  4. 4 CellCarta
  5. 5 Vector Institute
Correspondence: edolatab@yorku.ca