What if LLMs, SLMs and AI agents could search in biological embedding space instead of text?
Scientific discovery would get a lot better.
We are building a bulk transcriptomics foundation model that embeds 100M+ public RNA-seq samples spanning 20 years of research, drawn from comprehensive global nucleotide archives, processed functional genomics repositories, and specialized consortium portals.
Today, LLMs, SLMs, and AI agents search titles, abstracts, metadata, and keywords — but not the biology itself. Our model maps RNA-seq samples into a shared biological embedding space, enabling them to search by biological similarity rather than text.
This lets LLMs, SLMs, and AI agents connect experiments across diseases, tissues, and perturbations even when papers share no keywords. We ran a case study that surfaced novel hypotheses.
If you're harmonizing large RNA-seq datasets, we'd love to collaborate.
LLMs, SLMs (small language models), and AI agents can absolutely navigate and search natively within biological embedding spaces rather than relying strictly on human-written text. Moving beyond keyword-based or text-based literature review represents a major paradigm shift in how AI-driven scientific discovery operates — when these models navigate multi-dimensional mathematical representations of biology, such as cellular states, transcriptomic signatures, or protein structures, it fundamentally transforms their search capabilities.
Breaking the text bottleneck
Traditional scientific search engines, LLMs, SLMs, and AI agents rely on indexing text — titles, abstracts, and metadata. This means a model can only find what a human researcher explicitly chose to write down. If two different studies use entirely different terminology or look at different diseases, a text-based LLM or SLM might completely miss the hidden underlying connection.
By bypassing text and mapping data directly into a unified biological embedding space — such as RNA-seq transcriptomic signatures or protein-target latent spaces — the model can evaluate data based on literal molecular state and functional similarity.
Multimodal Biological Embeddings
Moving from keyword-based literature search — titles, abstracts, metadata — to high-dimensional continuous embedding spaces lets models capture the regulatory state and functional similarity of cells, tissues, and diseases directly from raw omics data.
Cross-Context Discovery
Projecting millions of public RNA-seq samples spanning decades of research into a shared latent space lets LLMs, SLMs, and AI agents identify hidden connections and shared mechanisms across entirely different diseases or perturbations that share zero overlapping keywords.
Patient Heterogeneity & Clinical Trials
Foundation models trained on bulk transcriptomics are increasingly applied to model complex patient heterogeneity, helping optimize patient stratification and translational decision-making for clinical trials.
LLM, SLM & Agentic Workflows
Integrating omics foundation models with advanced LLM and SLM reasoning frameworks — such as Claude Science or specialized multi-agent architectures — enables AI agents to run autonomous hypothesis generation, automated literature synthesis, and deeper mechanistic discovery.
How biological embedding search operates
Transcriptomic Signature Matching
Our platform maps diverse RNA-seq samples into a shared space. Instead of a naive search that simply matches the same tissues or diseases, advanced LLMs, SLMs, and AI agents look for geometric alignment — shared biological programs or patterns of variation across completely unrelated contexts.
Bypassing Human Bias
An LLM or SLM searching via embeddings can connect a rare disease study to an oncology trial if their molecular pathways react identically to a specific perturbation. Because the model evaluates raw data geometry rather than human keywords, it uncovers non-obvious hypotheses often hidden by academic silos.
Interpretable Latent Trees
Recent frameworks leverage regularized hyperbolic embeddings to visually map drug-target hierarchies. LLMs and SLMs use these structured spaces to predict unknown drug-protein interactions smoothly, without getting bogged down by irregular medical nomenclature.
Iterative Navigation & Goal-Setting
Cognition — synthetic and biological alike — can be modeled as the active remapping and navigation of embedding spaces to minimize error. Autonomous AI agents use internal memory loops to continuously update their position, steering toward target phenotypic profiles or optimized molecular structures.
The vision: an LLM, SLM & AI agent-powered web for labs
In the near future, specialized biological repositories won't just offer human-readable frontends. They will feature an LLM, SLM, and AI-agent-native layer where massive multi-omics data blocks are pre-translated into standardized embedding vectors.
This will allow ecosystems of LLMs, SLMs, and AI agents to fluidly pass complex biological concepts, cell state maps, and target structures between one another instantly — operating at a level of speed and mathematical precision that human language simply cannot achieve.
Built for the clinic, not just the bench
The model is designed to make bulk RNA-seq accessible in clinical settings, providing rich biological signals to capture a patient's disease state. The focus is on how it can learn patient heterogeneity to improve clinical trial decisions.
Bulk transcriptomics is a widely used omics modality across cell lines, animal models, and humans, making it a valuable tool for biological research. Future content will cover what the model learns and its applications in discovery, translation, and clinical contexts, including integration with Claude Science for hypothesis generation.
Hypothesis generation via biological embedding spaces
This represents a radical departure from traditional text-based AI scientific discovery. While text-based models — like standard LLMs — generate hypotheses by finding missing semantic links in published literature, geometry-driven biological embedding search bypasses human vocabulary entirely. It maps raw molecular, cellular, and structural data into multi-dimensional mathematical spaces, finding hidden causal relationships based on the actual rules of physics and biology.
The mechanics, frameworks, and core distinctions of this methodology define how it fundamentally accelerates discovery.
Text search vs. biological embedding search
| Feature | Text / Literature-Based Search | Biological Embedding Space Search |
|---|---|---|
| Data Input | Paper titles, abstracts, patents, metadata. | RNA-seq signatures, protein language models (pLMs), chemical vectors. |
| Core Mechanism | Token overlap, keyword co-occurrence, semantic text similarity. | Spatial proximity, geometric alignment, latent manifold trajectories. |
| Primary Bias | Human / reporting bias. Can only discover what humans choose to write about. | Measurement / data bias. Bound by the limitations and noise of wet-lab assays. |
| Discovery Type | Conceptual recombination — connects disconnected fields (e.g., neuroscience and bioelectricity). | Functional convergence — connects visually or textually distinct entities with identical molecular phenotypes. |
The three modes of embedding-driven hypothesis generation
1. Concept search vs. naive metric search
When searching standard cell data, a naive embedding search simply clusters data by "context neighbors" — pulling up the same tissue, species, or cell type because their baseline vectors look similar. True hypothesis generation uses concept search instead: rather than looking at absolute positions, the system queries the geometry of variation along a specific biological contrast. This finds an experiment where a completely different tissue or species shares a highly specific, obscure biological program.
2. Latent manifold trajectories
Biological processes follow low-dimensional pathways embedded in high-dimensional data — a principle known as the manifold hypothesis.
- A vector can represent a healthy cell state and a diseased cell state.
- The vector trajectory between them models disease progression.
- To generate therapeutic hypotheses, the system searches for a chemical or genetic perturbation vector that mathematically inverses that disease trajectory, pushing the cell back toward the healthy manifold.
3. Zero-shot evolutionary inferences
Protein language models (pLMs) like ProtT5 map protein sequences into structural embedding spaces. Searching this space doesn't require a paper explicitly stating that two proteins are related — by visualizing and analyzing the distances between out-of-distribution sequences, this approach can identify evolutionary convergence, such as tracking how secreted toxin families evolved from membrane-anchored proteins, generating entirely new structural biology hypotheses.
State-of-the-art frameworks (2026)
Universal Cell Embedding (UCE)
Published in Nature (July 2026), UCE acts as a foundation model for single-cell biology. It embeds tens of millions of cells across multiple species into a unified atlas, exhibiting emergent behavior — accurately predicting developmental lineages and tissue organization it was never explicitly trained on — making it a premier tool for zero-shot cell-state hypothesis generation.
Graph-PRefLexOR
This framework combines neural generation with graph-native reinforcement learning. It explores bounded biological and materials embedding spaces, enforcing strict structural reasoning and traceability — providing a 2× to 3× increase in semantic diversity and conceptual recombination compared to standard text models.
ProtSpace
A specialized engine that bridges structural biology and machine learning embeddings, allowing interactive mapping, visualization, and analysis of structural protein spaces to refute or validate evolutionary and functional hypotheses.
The core bottleneck: evaluation
The primary obstacle in shifting from text to embedding search is evaluating hypothesis quality. Text models can be scored against historic literature or cross-validated by human reading. Embedding-generated hypotheses often predict interactions or mechanisms that have never been observed or documented by humans. Because retrospective benchmarks are heavily biased toward what is already known, evaluating these deep latent discoveries ultimately relies on real-world closed-loop wet-lab validation and active scientist feedback loops.
Building an architecture for this?
We'd love to hear from you — whether you're harmonizing large RNA-seq datasets, evaluating pre-trained biology foundation models like UCE or ESM versus a custom contrastive embedding space, or deciding between unsupervised clustering and supervised link-prediction over a hybrid knowledge graph. We're happy to go deeper on the loss functions and retrieval pipelines behind this work.