Understanding Tox alleles in BIGSdb
The presence or absence of a toxin gene in a BIGSdb isolate record should be interpreted with caution.
It is important to distinguish between several different levels of information:
- the presence of a tox gene sequence in the assembly;
- the existence of a corresponding allele in the BIGSdb database;
- the toxigenic prediction associated with this allele;
- and finally, the actual expression and functionality of the toxin by the strain.
These different levels do not necessarily correspond to the same biological reality.
🔍 How is a tox allele identified?
For a tox allele to be associated with an isolate in BIGSdb, the sequence found in the genome must match exactly one of the reference alleles already described in the allele nomenclature database.
The match must have:
- 100% identity: no nucleotide differences compared with the reference allele;
- 100% coverage: the entire allele sequence must be present.
This approach allows allele numbers to be assigned in a reliable and reproducible way. However, it also means that the absence of a tox allele from an isolate record does not necessarily mean that the tox gene is absent from the genome.
⚠️ Tox allele not found
Important: the absence of a tox allele in an isolate record does not necessarily mean that the tox gene is absent.
If a toxin allele is not found by BIGSdb, several situations are possible.
🧩 Assembly issue
The toxin gene may be present in the strain but not correctly represented in the genome assembly.
For example:
- the gene may be located in a region that could not be assembled and is therefore absent from the available contigs;
- the gene may be located at the end of a contig and therefore only be partially assembled;
- the sequence may be fragmented across several contigs.
In these situations, the sequence available in the assembly does not allow BIGSdb to identify the complete allele.
The toxin may therefore be present in the strain without being detectable in the assembly.
🧬 The gene is present in the assembly but sufficiently different
Another possibility is that the toxin gene is present in the assembly, but its sequence differs from the alleles currently available in the database.
For example, the sequence may contain:
- an insertion sequence (IS);
- recombination events;
- highly variable regions within the toxin gene;
- or other sequence changes that prevent a perfect match with a known allele.
In this situation, the sequence may clearly correspond to a toxin gene, but it does not meet the criteria required to be assigned an existing allele number.
If no new allele has yet been defined for this sequence, it will therefore not appear as a tox allele in the isolate record.
⏳ Curation delay
A third possibility is that the sequence represents a new tox allele that has not yet been curated in the database.
If a toxin sequence differs from all existing alleles, a new allele may need to be defined and added to BIGSdb before it can be assigned to the corresponding isolates.
Therefore, there may be a delay between:
- the identification of a new or divergent tox sequence in an assembly;
- the curation and validation of the sequence;
- the creation of a new allele in BIGSdb;
- and its subsequent assignment to the relevant isolates.
During this period, an isolate may appear as tox not found, even though a tox gene sequence is present in its genome assembly.
The time between the completion of the automated curation process (definition of new STs, cgMLSTs, etc.) and the point at which the new tox allele is defined and then assigned to your isolate can take anywhere from several hours to several days.
Generally, results are available within one week; please do remember to check the details of your isolates yourself.
💻 tox scheme and toxigenic prediction
The tox scheme and the associated toxigenic prediction are based on the genomic information available.
They therefore represent a prediction based on the DNA sequence, rather than a direct measurement of the biological activity of the toxin.
The presence of an allele predicted to be toxigenic means that the sequence meets the genomic criteria used for this prediction. It does not, by itself, demonstrate that the toxin is actually produced and functional in the strain.
A protein predicted to be toxigenic may not be functional
A sequence may correspond to a toxin predicted to be functional while the resulting protein is, in reality, non-functional.
For example, a single-nucleotide polymorphism (SNP) may change an amino acid that is important for the structure or function of the protein.
Even if the sequence still meets the criteria for a toxigenic prediction, such a mutation may:
- alter the three-dimensional structure of the protein;
- disrupt its active site;
- or prevent the protein from interacting correctly with its target.
Therefore, a genomic prediction (even when Elek’s test is positive for isolates with this particular allele) does not allow to predict the actual functionality of the resulting protein.
A toxigenic protein-coding gene may be present without the strain expressing the toxin
Conversely, a strain may contain a gene encoding a functional toxigenic protein gene, without actually producing the toxin.
Gene expression depends not only on the coding sequence, but also on regulatory regions and other mechanisms controlling transcription.
For example, a mutation or insertion in the promoter region may prevent or substantially reduce transcription of the toxin gene. In this situation:
- the gene may encode a functional toxigenic protein;
- the sequence may therefore be correctly identified in BIGSdb’s database;
- but the strain may not produce the toxin, or may produce it at a greatly reduced level.
Example: tox 19, IS1132 and toxin expression
Isolates 366 and 367 provide an example of the difference between a genomic toxigenic prediction and the actual production of the toxin.
Both isolates carry the tox-19 allele, which is predicted to be toxigenic based on its sequence. However, both isolates are Elek-test negative, indicating that the toxin is not detected as being produced by these strains.
Genomic analysis provides a possible explanation for this discrepancy: an IS1132 insertion is present upstream of the tox-19 gene, in a region involved in the regulation of gene expression. This insertion disrupts the promoter or otherwise interferes with transcription of the toxin gene.
This example demonstrates that:
The presence of a toxigenic tox allele does not necessarily mean that the strain can express the diphtheria toxin.
The coding sequence and its regulatory context must therefore be considered separately when interpreting the biological significance of a tox prediction.
🔑 Key points
The interpretation of tox data in BIGSdb should therefore be considered at several levels:
-
tox allele identified → The sequence present in the assembly matches an allele in the database with 100% identity and 100% coverage.
-
tox not found → This does not prove that the gene is absent. The gene may be missing from the assembly, partially assembled, sufficiently different from known alleles, or represented by a new allele that has not yet been curated.
-
Positive toxigenic prediction → This is a prediction based on genomic sequence. It is not direct evidence that the toxin is produced or biologically active.
-
Functionality and expression → These depend on the protein sequence as well as other factors, including regulatory regions and mechanisms controlling gene expression.
-
An Elek positive isolate may produce a non-functional diphtheria toxin. Only the confirmation of toxigenicity e.g. in in-vitro or in-vivo assays, or association with toxinic syndrome in clinical data, can confirm toxigenicity.
In summary, BIGSdb primarily describes what can be inferred from the available genomic sequence. They should not automatically be interpreted as a direct measurement of the biological phenotype of the strain.
Experimental confirmation is required to determine with certainty whether a toxin is actually expressed (Elek test) and functional (in-vivo or in-vitro phenotypes).