프로젝트 기록

PROJECT RECORD

UNK-g: narrowing the identity of an unassigned protein

한국어로 읽기

A protein structure can show the shape of a component without telling us which gene produces it. In the public records examined here, one component still carries a placeholder rather than an assigned identity: UNK-g.

It appears in structural records concerning mitochondrial ribosome assembly in Trypanosoma brucei. A ribosome reads genetic information to make proteins; mitochondria have their own version of this machinery. In this case, a component seen in an assembly structure has not been linked to a gene. I brought together the public structural records and findings from the existing analysis to explain which candidate they support and where that support stops.

A candidate among 7,266 predicted models

In the underlying computational comparison, Q38DP6 ranked first among 7,266 retained predicted protein models. These models estimate the shapes proteins might take, making them candidates to compare with the unassigned shape. The gene identifier associated with Q38DP6 is Tb927.9.11350.

The analysis compared records 6SG9, 9HNY and 7PUA, finding compatibility with the unassigned component in more than one deposited model. This is not three independent confirmations. In particular, 9HNY reinterprets previously collected data, and several assessments reuse the same structural observations. Agreement between them adds correlated support: different views of shared evidence, rather than entirely separate observations.

Does the shape fit what was actually observed?

Resembling a structural model and fitting the observations behind it are different checks. A density map contains spatial information from the observations and constrains where a protein model can fit. Trying a part against an outline is a useful analogy, but the experimental outline is itself uncertain; a plausible visual match cannot settle the identity.

On a density-fit measure evaluated in a fixed region, Q38DP6 was favored over seven alternative candidates. Yet after refinement, it fitted the local density less well than the original unassigned trace. Being better than the other candidates is not the same as matching well enough to replace the original model. This comparison supports the nomination without establishing the assignment.

Published measurements of proteins in samples provide a different kind of evidence. The underlying report describes the candidate appearing more abundantly in an assembly-associated dataset and showing a similar migration pattern to assembly-related proteins in a separate analysis. This co-migration is a pattern in the analysis, rather than direct observation of the proteins moving together inside the organism. The observations support an association with ribosome assembly, while leaving direct binding and the identity of this density unresolved.

First place is not a probability of being right

A firm assignment would need more than an overall resemblance. The local structural information here does not allow the amino-acid sequence to be read unambiguously. An earlier sequence-assessment score lost discriminatory power after refinement. A later classifier, a computational tool for distinguishing candidates, also favored this candidate. Its accuracy was limited on held-out data: data that had not been used to build the tool.

To interpret a score, we also need to know how it behaves for incorrect or chance matches. That reference distribution was not adequately calibrated in this analysis. I therefore do not turn the reported significance measures or first-place ranking into a probability that the identity is correct. Density beyond the original trace is weak, so apparent coverage at a permissive threshold carries limited weight. An unfinished automated modeling attempt contributes no evidence.

What I have released

I published this as a preliminary, non-peer-reviewed research summary on 8 October 2026. It narrows the identity question to a named candidate, Q38DP6; it does not report a correction accepted by the original data providers. UNK-g’s identity and biochemical role remain unresolved, and the analysis establishes neither a direct interaction nor a therapeutic application.

The summary selects findings and limitations from the existing analysis. It contains no sequences, coordinates, executable code or experimental procedures, and is not a complete reproducibility package. The source datasets remain attributable to their original investigators. No new biological calculation was performed for this edition. AI assistance supported comparison of statements across the supplied reports, identification of reporting inconsistencies, and drafting and editing of the summary.

Research summary and references — DOI 10.5281/zenodo.23233834

연구 목록 · 작업 목록

Research index · Work index

Discover more from Woong Works

Subscribe now to keep reading and get access to the full archive.

Continue reading