Nour El Kazwini, Mingze Gao, Idris Kouadri Boudjelthia, Fangxin Cai, Yuanhua Huang, Guido Sanguinetti
7 min
RNA velocity has recently emerged as a key tool in the analysis of single-cell transcriptomic data, yet connecting RNA velocity analyses to underlying regulatory processes has proved challenging. Here, we propose CRAK-Velo, a semi-mechanistic model that integrates chromatin accessibility data in the estimation of RNA velocities. CRAK-Velo provides biologically consistent estimates of developmental flows and enables accurate cell-type deconvolution, while additionally shining light on regulatory processes at the level of interactions between genes and chromatin regions.
RNA velocity is a powerful tool for inferring developmental trajectories from single-cell transcriptomic data, but it often struggles to connect observed RNA dynamics to the underlying regulatory mechanisms. Existing methods either rely solely on splicing information or require complex parameterization to incorporate epigenomic data. The authors introduce CRAK-Velo to address these limitations by explicitly integrating chromatin accessibility data to regularize and improve the estimation of gene transcription rates.
CRAK-Velo builds upon the parametric framework of UniTVelo but introduces a mechanism to link transcription rates directly to chromatin accessibility. The model uses a probabilistic topic model (cisTopic) to smooth sparse scATAC-seq data and compute open chromatin probabilities for regions near genes. By incorporating these accessibility scores into a likelihood function, the model reconciles data-driven splicing kinetics with chromatin-based production rates. This allows the researchers to quantify the specific contribution of individual regulatory regions to transcriptional dynamics over pseudotime.
When tested on hematopoietic stem cell differentiation and mouse embryonic brain development datasets, CRAK-Velo demonstrated superior performance in reconstructing developmental flows compared to existing methods like UniTVelo and MultiVelo. Specifically, it correctly identified terminal cell states that other models misclassified as intermediate stages. Furthermore, the model provided more accurate cell-type deconvolution, as evidenced by higher classification accuracy using the inferred chromatin-unspliced read representation. The authors also demonstrated the model's utility in visualizing the regulatory dynamics of key genes, such as KLF1 and MSI2, by mapping the influence of proximal chromatin regions over time.
By providing a more biologically consistent way to integrate multi-omic data, CRAK-Velo offers a clearer window into the regulatory logic of gene expression. Its ability to quantify the impact of specific chromatin regions on transcription makes it a valuable tool for generating testable hypotheses regarding gene regulation in complex developmental processes.
Sam: And those patterns become the "normal" reference for what the cell's regulatory landscape should look like at each stage of development?
Alex: Exactly. It's like building a reference library of what healthy, plausible regulatory states look like. If the model predicts a cell trajectory that doesn't match any state in that library, it gets flagged as implausible. It acts as a biological safety fence.
Sam: And the name CRAK-Velo — what does that stand for?
Alex: It stands for Chromatin Accessibility Kinetics integration in RNA Velocity. The "kinetics" part is important — it's not just asking whether chromatin is open, but how quickly it's opening or closing, and whether that rate matches what the RNA output would predict. It's a semi-mechanistic model, meaning it combines biological rules with data-driven learning. It's not a pure black box.
Sam: So it's constrained by the physical reality of how genes are actually switched on, rather than just finding statistical patterns in the data.
Alex: That's the key distinction. And that constraint actually makes the model more efficient computationally, because it narrows down the range of plausible answers the algorithm has to search through.
Sam: Does this actually change the results when you test it on real biological data?
Alex: It does, in a meaningful way. In a test using human stem cells — cells that can develop into many different types — the model correctly identified three separate terminal cell states. Terminal states are the final destinations: the mature cell types a stem cell can become. Other methods got confused and predicted impossible flows between those states, essentially drawing roads through mountains. CRAK-Velo drew the map correctly.
Sam: So it's not just a theoretical improvement. It changes the biological conclusions you'd actually draw from the data.
Alex: Right. And it also performs better at distinguishing different cell types within a mixed population — which matters a great deal when you're studying something like a developing organ, where many cell types are present at once.
Sam: You mentioned it assumes that RNA production is linked to the accessibility of nearby regulatory regions. How nearby are we talking?
Alex: The model looks within a window of ten thousand base pairs upstream or downstream from where a gene starts. That's a relatively short distance on the scale of the genome. It focuses on what are called proximal regulatory regions — the local switches right next to the gene.
Sam: What if the important switch is further away?
Alex: That's a genuine limitation the authors acknowledge. Long-range regulatory interactions — where a switch far from the gene controls its activity — are outside the model's scope. It's a trade-off. By focusing on the local region, the model stays tractable and interpretable, but it may miss some real biological signals.
Sam: Does the model require specially collected data, or can it work with existing datasets?
Alex: It requires what's called Multiome data — a relatively recent technology that captures both RNA activity and chromatin accessibility from the exact same single cell at the same time. That pairing is essential, because you need both signals to be matched. The limitation is that this type of data is still less common than standard single-cell RNA data, so the model can't be applied to older datasets that only captured one signal.
Sam: And the authors mention other limitations beyond the window size?
Alex: Yes. The model still struggles to identify certain specific terminal cell states — ependymal cells, which line the fluid-filled cavities of the brain, were one example where it fell short. And like any model that combines two data types, it's sensitive to the quality of how those two datasets are aligned. If the alignment is poor, the predictions suffer.
Sam: Where do the authors see this going from here?
Alex: The paper suggests future versions could incorporate spatial transcriptomics — which tells you not just what genes are active, but where in a tissue the cell sits physically. Or protein-level data, which would add yet another layer of biological reality. The goal would be to map not just what a cell is doing, but the precise physical environment that's triggering those changes.
Sam: So the broader ambition is to build a model of development that reflects more of the actual biology, layer by layer.
Alex: That's a fair summary. By linking the chromatin state — the throttle — to the RNA output — the speedometer — we get a clearer and more honest picture of what a cell is actually doing and where it's actually headed. And that kind of accuracy matters when you're trying to understand how tissues form, how diseases develop, or how to guide cells toward a particular fate in a therapeutic context.
Sam: Thanks for walking me through the mechanics of this, Alex. It's a good reminder that sometimes the most useful thing you can add to a model is a constraint — a rule that says, "this path isn't allowed."
Alex: Well put. Adding biological constraints doesn't limit what the model can find — it helps it find things that are actually true. Thanks for listening to ResearchPod.