DeepSomatic was trained on three of the breast cancer genomes and the two lung cancer genomes in the CASTLE reference dataset, then tested on data it had not seen during training, including the single breast cancer genome that was excluded and chromosome 1 from each sample.
In those tests, DeepSomatic models developed for each of the three major sequencing platforms performed better than other methods, identifying more tumor variants with higher accuracy. For short-read sequencing data, the comparison tools were SomaticSniper, MuTect2 and Strelka2 (with SomaticSniper specifically for single nucleotide variants, or SNVs). For long-read sequencing data, DeepSomatic was compared against ClairS, a deep learning model trained on synthetic data.
Across the six reference cell lines and a seventh preserved sample, DeepSomatic identified 329,011 somatic variants. The model performed particularly well on cancer variations involving insertions and deletions, or “Indels”, improving the F1-score, which measures how well a model finds true variants while avoiding false positives.
On Illumina sequencing data, the next-best method scored 80% at identifying Indels, while DeepSomatic scored 90%. On Pacific Biosciences sequencing data, the next-best method scored less than 50% at identifying Indels, and DeepSomatic scored more than 80%.
Source: research.google.
Companies can share verified announcements through Newz9’s international press release submission page.

