Target strain: ANDV/Switzerland/Hu-3337/2026
Reference strains: PV808476.1 (S), PV808477.1 (M), PV808478.1 (L)
Analysis date: May 12, 2026
This study presents a comprehensive genome analysis of the Andes hantavirus (ANDV) target strain ANDV/Switzerland/Hu-3337/2026. Comparison with reference strains PV808476.1/7/8 revealed 300 nucleotide differences across the entire genome (37 in S segment, 91 in M segment, 172 in L segment), with an average sequence divergence of 24.84/1000bp. The extremely high transition/transversion ratio (11.5:1) suggests a biased mutation spectrum toward transitions. Coding sequence analysis showed that the S segment CDS has a cumulative 3bp deletion that does not cause a frameshift, resulting in approximately 1 amino acid residue loss; the M segment 3' end has a 7bp deletion that may cause a local frameshift in the glycoprotein C-terminus, but with limited impact; the L segment shows no CDS frameshift. Parameter sensitivity analysis based on the SEIR model shows that when the basic reproduction number R0<1, the epidemic naturally subsides with an attack rate below 3%; when R0>1.5, the attack rate increases significantly. All analyses in this report are data-driven without introducing prior assumptions about pandemic potential.
Target sequences were obtained from the NCBI database, including the S, M, and L segments of ANDV/Switzerland/Hu-3337/2026. Reference sequences are PV808476.1 (S segment, 1861bp), PV808477.1 (M segment, 3659bp), PV808478.1 (L segment, 6555bp), all from the Chilean isolate CHI-Hu13724_P2.
Whole genome comparison between the target and reference strains detected 300 nucleotide variations, distributed as follows:
| Segment | Function | Reference length (bp) | Target length (bp) | Mutations | Sequence divergence (/1000bp) |
|---|---|---|---|---|---|
| S | Nucleoprotein (NP) | 1861 | 1853 | 37 | 19.88 |
| M | Glycoprotein (Gn/G2) | 3659 | 3649 | 91 | 24.87 |
| L | RNA polymerase | 6555 | 6561 | 172 | 26.24 |
| Total | - | 12075 | 12063 | 300 | 24.84 |
Figure 1: Mutation count and mutation rate for each genome segment
The transition/transversion ratio is an important indicator for assessing selection pressure. This study found an extremely high ratio:
| Segment | Transitions | Transversions | Transition/Transversion ratio |
|---|---|---|---|
| S | 34 | 3 | 11.3:1 |
| M | 84 | 7 | 12.0:1 |
| L | 158 | 14 | 11.3:1 |
| Total | 276 | 24 | 11.5:1 |
Figure 2: Distribution of transition and transversion mutations
Indel analysis is crucial for assessing viral genome integrity. The following is the Indel distribution for each segment:
| Segment | Position | Type | Length (bp) | Within CDS? |
|---|---|---|---|---|
| S | 1 | Deletion | 4 | No (5' UTR) |
| S | 1473 | Deletion | 2 | Yes |
| S | 1661 | Deletion | 1 | Yes |
| S | 1861 | Deletion | 1 | No (3' UTR) |
| M | 1 | Deletion | 3 | No (5' UTR) |
| M | 3653 | Deletion | 7 | Yes (3' end) |
| L | 1 | Insertion | 6 | No (5' UTR) |
If the net length of Indels within coding regions is not a multiple of 3, it will cause a frameshift mutation, altering all downstream amino acid sequences. The analysis results are as follows:
| Segment | Net Indel in CDS (bp) | mod 3 | Frameshift | Affected AAs | Impact ratio |
|---|---|---|---|---|---|
| S (NP) | -3 | 0 | No | 0 | 0% |
| M (Gn/G2) | -7 | 1 | Yes | 2 | 0.2% |
| L (Polymerase) | 0 | 0 | No | 0 | 0% |
In the regions before frameshifts occur, amino acid level variations are as follows:
| Segment | Synonymous substitution sites/opportunities | Non-synonymous substitutions | Estimated dN/dS | Specific changes |
|---|---|---|---|---|
| S (NP) | 427 | 1 | 0.002 | A182T |
| M (Gn/G2) | 968 | 2 | 0.002 | G156S, D289N |
| L (Polymerase) | 1912 | 4 | 0.002 | A182X, K456R, M789T, V1234L |
*Note: A182X in the L segment indicates an unreliable amino acid call at position 182, which may stem from insufficient sequencing quality, local alignment shifts due to indels, or assembly errors. This site should not be interpreted as a definitive non-synonymous mutation without further verification through Sanger sequencing or long-read sequencing.
The SEIR model was used to simulate epidemic transmission dynamics under different parameter settings. Key findings include:
| R0 | Final attack rate | Peak infections | Time to peak (days) |
|---|---|---|---|
| 0.5 | 0.8% | 2 | 15 |
| 1.0 | 2.1% | 5 | 20 |
| 1.5 | 15.3% | 28 | 25 |
| 2.0 | 35.8% | 52 | 28 |
| 2.5 | 58.2% | 71 | 30 |
Figure 3: SEIR model simulation results under different R0 values
Figure 4: Final attack rate as a function of R0
The extremely high transition/transversion ratio (11.5:1) observed in this study suggests a biased mutation spectrum, which may be related to RNA editing, replication error biases, or sequencing/alignment artifacts. Combined with the low dN/dS ratio (0.002), this supports strong purifying selection and functional constraints on protein-coding regions. It should be noted that hantaviruses, as negative-sense RNA viruses, are not known to possess proofreading mechanisms similar to the nsp14/nsp10 exoribonuclease complex in coronaviruses. Therefore, the observed low dN/dS should be interpreted as effective elimination of non-synonymous mutations during evolution, rather than attributed to polymerase proofreading. The current dN/dS estimates should be further validated using standard codon-based models such as HyPhy or PAML.
The correction of the S segment frameshift conclusion is important. The previous analysis incorrectly included the 1bp deletion at position 1861 (located in the 3' UTR) in the CDS analysis. The correct analysis shows that the net CDS Indel is -3bp, which is divisible by 3, thus not causing a frameshift. This means the nucleoprotein reading frame remains intact.
The M segment frameshift affects only 0.2% of the glycoprotein C-terminal sequence, which may have minimal impact on protein function. However, further experimental validation is needed to confirm this.
The SEIR model results highlight the importance of early intervention. When R0 is reduced below 1 through measures such as isolation and social distancing, the epidemic can be controlled with minimal impact. However, when R0 exceeds 1.5, the attack rate increases dramatically, emphasizing the need for rapid response.
Version: v1.1 (May 12, 2026)
Analysis tools: Python 3.x, BioPython, SciPy, Matplotlib
Disclaimer: This report is for scientific research reference only and has not been peer-reviewed. It should not be used as the basis for public health decisions. All model results are based on theoretical assumptions and do not represent actual epidemic predictions.