Carbon-A open model and database predict 566 million gene candidates across 22,617 species
Original titleWe're releasing Carbon-A: an open model to discover new genes in DNA.
AISummary
Carbon-A is an open model that predicts gene locations directly from DNA, and it has been used to annotate genomes from over 22,000 species. The release includes a database of 566 million gene candidates, about 16 times the gene annotations in the RefSeq dataset. Wet-lab RNA experiments supported 239 candidates missing from RefSeq across cats, Syrian hamsters, chickens, and Arabidopsis.
AIWhy it matters
The source ties an open gene-annotation model to specific wet-lab checks and gene counts, helping readers judge how far its predictions extend beyond well-studied genomes.
Source: Leandro von Werra · x.com