Reading the letters of genetic information correctly does not guarantee that we have interpreted them correctly. Proteins are built by joining smaller components called amino acids. Genetic information specifies their order, and the rules that connect combinations of letters to amino acids are called the genetic code. A dictionary offers a useful comparison: knowing how a word is spelled is different from knowing what it means. In biology, the complication is that not every organism uses exactly the same code.
The question in this research summary concerns mitochondrial records from the fungus Pneumocystis jirovecii. Mitochondria are structures inside cells that carry genetic information of their own. At some positions, the current interpretation of the public records disagrees with protein features that are conserved across related organisms. One candidate explanation is that the decoding rule differs from the rule used in the current records.
Agreement across records is not the same as independent confirmation
The source analysis examined 24 public records of P. jirovecii and found a recurring pattern. Those 24 records are not 24 independent biological samples. The set includes a curated reference and the source record behind it, as well as partial records. Keeping both an original document and an edited copy gives us two files, but not two independent sources of information.
In preparing the summary, I kept consistency across records separate from confirmation across independent samples. The recurrence gives us a reason to examine the hypothesis, but the effective number of independent biological observations is not established by this record count. Shared evolutionary history and reuse of comparative evidence also limit how independently the observations can support the conclusion.
What RNA can tell us, and what it cannot
RNA is an intermediate between genetic information and the production of proteins. One competing explanation is that the information changes at this intermediate stage. The public RNA dataset instead supports retention of the information under examination, weakening that particular explanation within this dataset. This still does not identify the composition of the final protein. Checking a copy of an instruction sheet is different from checking the finished object: the RNA observation does not establish which amino acid was incorporated into protein.
The findings are therefore indirect. They leave the current interpretation in contention, and the analysis provides neither direct protein-level evidence that distinguishes the alternatives nor a demonstrated mechanism for the proposed decoding rule.
Counting analysis runs without turning them into samples
The summary also distinguishes the reported analysis totals. The source report describes 4,775 runs, while the consolidated result table contains 4,743 data rows. The remaining 32 runs are described as follow-up analyses. The larger number is a reported total across stages, not the row count of that table, not a count of independent biological samples, and not a claim that every run has been independently reproduced.
For the RNA dataset, the aggregate figure is 3,886 reads supporting retention of the information under examination out of 3,891 total reads. A read is a small piece of data obtained by reading RNA information. Many pieces from the same dataset do not become independent samples, or evidence of the final protein, simply because there are many of them.
What the public summary leaves unresolved
Not finding a usable public dataset limits the evidence available to this analysis; it does not prove that such data do not exist. Similarly, not locating an earlier report does not establish novelty or priority.
I released the two-page research summary on 8 October 2026, separating the candidate interpretation from what the existing analyses and their numerical totals support. It is preliminary and not peer reviewed, and the evidence is insufficient to replace the current annotation. The summary makes no diagnostic or therapeutic claim and is not a complete reproducibility package. Its contribution is a candidate explanation for a recurring discrepancy, with the unresolved alternatives left visible.
Read the public research summary and references ↗ · DOI: 10.5281/zenodo.23233894
