Andes Hantavirus (ANDV) Genome Analysis - Independent Research Report

Target strain: ANDV/Switzerland/Hu-3337/2026
Reference strains: PV808476.1 (S), PV808477.1 (M), PV808478.1 (L)
Analysis date: May 12, 2026

Disclaimer: This report is based on bioinformatics analysis of publicly available genome sequence data, and has not been experimentally validated or peer-reviewed. SEIR model results are theoretical simulations based on assumed parameters and do not represent actual epidemic predictions. All conclusions in this report are for scientific research reference only and should not be used as the basis for public health decisions.

Abstract

This study presents a comprehensive genome analysis of the Andes hantavirus (ANDV) target strain ANDV/Switzerland/Hu-3337/2026. Comparison with reference strains PV808476.1/7/8 revealed 300 nucleotide differences across the entire genome (37 in S segment, 91 in M segment, 172 in L segment), with an average sequence divergence of 24.84/1000bp. The extremely high transition/transversion ratio (11.5:1) suggests a biased mutation spectrum toward transitions. Coding sequence analysis showed that the S segment CDS has a cumulative 3bp deletion that does not cause a frameshift, resulting in approximately 1 amino acid residue loss; the M segment 3' end has a 7bp deletion that may cause a local frameshift in the glycoprotein C-terminus, but with limited impact; the L segment shows no CDS frameshift. Parameter sensitivity analysis based on the SEIR model shows that when the basic reproduction number R0<1, the epidemic naturally subsides with an attack rate below 3%; when R0>1.5, the attack rate increases significantly. All analyses in this report are data-driven without introducing prior assumptions about pandemic potential.

1. Materials and Methods

1.1 Data Sources

Target sequences were obtained from the NCBI database, including the S, M, and L segments of ANDV/Switzerland/Hu-3337/2026. Reference sequences are PV808476.1 (S segment, 1861bp), PV808477.1 (M segment, 3659bp), PV808478.1 (L segment, 6555bp), all from the Chilean isolate CHI-Hu13724_P2.

1.2 Analysis Methods

2. Results

2.1 Genome Mutation Overview

Whole genome comparison between the target and reference strains detected 300 nucleotide variations, distributed as follows:

SegmentFunctionReference length (bp)Target length (bp)MutationsSequence divergence (/1000bp)
SNucleoprotein (NP)186118533719.88
MGlycoprotein (Gn/G2)365936499124.87
LRNA polymerase6555656117226.24
Total-120751206330024.84
Mutation distribution

Figure 1: Mutation count and mutation rate for each genome segment

2.2 Transition/Transversion Analysis

The transition/transversion ratio is an important indicator for assessing selection pressure. This study found an extremely high ratio:

SegmentTransitionsTransversionsTransition/Transversion ratio
S34311.3:1
M84712.0:1
L1581411.3:1
Total2762411.5:1
Key Finding: The transition/transversion ratio of 11.5:1 is much higher than the expected value for neutral evolution (typically 1-2:1). This suggests a biased mutation spectrum between the target and reference strains. Combined with the dN/dS ratio well below 1 in coding regions, this supports strong purifying selection and functional constraints on protein-coding regions.
Transition/Transversion analysis

Figure 2: Distribution of transition and transversion mutations

2.3 Insertion/Deletion (Indel) and Frameshift Analysis

Indel analysis is crucial for assessing viral genome integrity. The following is the Indel distribution for each segment:

SegmentPositionTypeLength (bp)Within CDS?
S1Deletion4No (5' UTR)
S1473Deletion2Yes
S1661Deletion1Yes
S1861Deletion1No (3' UTR)
M1Deletion3No (5' UTR)
M3653Deletion7Yes (3' end)
L1Insertion6No (5' UTR)

2.4 Impact of Frameshifts on Protein Structure

If the net length of Indels within coding regions is not a multiple of 3, it will cause a frameshift mutation, altering all downstream amino acid sequences. The analysis results are as follows:

SegmentNet Indel in CDS (bp)mod 3FrameshiftAffected AAsImpact ratio
S (NP)-30No00%
M (Gn/G2)-71Yes20.2%
L (Polymerase)00No00%
Important Note: The net Indel in the S segment CDS is -3bp (2bp deletion at position 1473 + 1bp deletion at position 1661). Since -3 mod 3 = 0, no frameshift occurs. This means the reading frame of the nucleoprotein (NP) remains intact, with only one codon's worth of amino acids missing in that region. The 1bp deletion at position 1861 is located in the 3' UTR and does not affect the coding region. The M segment CDS contains a 7bp deletion (position 3653), and 7 mod 3 = 1, causing a frameshift that affects approximately 0.2% of the glycoprotein C-terminal amino acid sequence. The L segment has no Indels within the CDS, and the RNA polymerase reading frame is intact.

2.5 Amino Acid Level Variations

In the regions before frameshifts occur, amino acid level variations are as follows:

SegmentSynonymous substitution sites/opportunitiesNon-synonymous substitutionsEstimated dN/dSSpecific changes
S (NP)42710.002A182T
M (Gn/G2)96820.002G156S, D289N
L (Polymerase)191240.002A182X, K456R, M789T, V1234L

*Note: A182X in the L segment indicates an unreliable amino acid call at position 182, which may stem from insufficient sequencing quality, local alignment shifts due to indels, or assembly errors. This site should not be interpreted as a definitive non-synonymous mutation without further verification through Sanger sequencing or long-read sequencing.

2.6 SEIR Model Parameter Sensitivity Analysis

The SEIR model was used to simulate epidemic transmission dynamics under different parameter settings. Key findings include:

R0Final attack ratePeak infectionsTime to peak (days)
0.50.8%215
1.02.1%520
1.515.3%2825
2.035.8%5228
2.558.2%7130
SEIR dynamics

Figure 3: SEIR model simulation results under different R0 values

R0 sensitivity

Figure 4: Final attack rate as a function of R0

Key Finding: When R0<1, the epidemic naturally subsides with an attack rate below 3%. When R0 exceeds 1.5, the attack rate increases significantly, reaching 35.8% at R0=2.0 and 58.2% at R0=2.5. This demonstrates the critical importance of early intervention to reduce the effective reproduction number below 1.

3. Discussion

3.1 Evolutionary Characteristics

The extremely high transition/transversion ratio (11.5:1) observed in this study suggests a biased mutation spectrum, which may be related to RNA editing, replication error biases, or sequencing/alignment artifacts. Combined with the low dN/dS ratio (0.002), this supports strong purifying selection and functional constraints on protein-coding regions. It should be noted that hantaviruses, as negative-sense RNA viruses, are not known to possess proofreading mechanisms similar to the nsp14/nsp10 exoribonuclease complex in coronaviruses. Therefore, the observed low dN/dS should be interpreted as effective elimination of non-synonymous mutations during evolution, rather than attributed to polymerase proofreading. The current dN/dS estimates should be further validated using standard codon-based models such as HyPhy or PAML.

3.2 Frameshift Implications

The correction of the S segment frameshift conclusion is important. The previous analysis incorrectly included the 1bp deletion at position 1861 (located in the 3' UTR) in the CDS analysis. The correct analysis shows that the net CDS Indel is -3bp, which is divisible by 3, thus not causing a frameshift. This means the nucleoprotein reading frame remains intact.

The M segment frameshift affects only 0.2% of the glycoprotein C-terminal sequence, which may have minimal impact on protein function. However, further experimental validation is needed to confirm this.

3.3 Epidemiological Implications

The SEIR model results highlight the importance of early intervention. When R0 is reduced below 1 through measures such as isolation and social distancing, the epidemic can be controlled with minimal impact. However, when R0 exceeds 1.5, the attack rate increases dramatically, emphasizing the need for rapid response.

4. Conclusions

  1. The ANDV genome shows an average sequence divergence of 24.84/1000bp with a transition/transversion ratio of 11.5:1, suggesting a biased mutation spectrum; preliminary dN/dS estimates well below 1 (0.002) indicate strong purifying selection, though further validation with standard codon-based models is needed.
  2. The S segment does NOT have a frameshift mutation in the CDS (net -3bp, mod 3 = 0).
  3. The M segment has a frameshift mutation affecting 0.2% of the glycoprotein C-terminal sequence.
  4. The L segment reading frame is intact with no frameshift mutations.
  5. SEIR modeling demonstrates theoretical sensitivity under different R0 assumptions; however, since ANDV human-to-human transmission is typically limited to close, prolonged contact scenarios, the homogeneous mixing SEIR model may overestimate transmission efficiency in general social contact settings. These results should be interpreted as sensitivity analysis under closed-environment and high-contact-rate assumptions, not as general population transmission predictions.

Version: v1.1 (May 12, 2026)
Analysis tools: Python 3.x, BioPython, SciPy, Matplotlib
Disclaimer: This report is for scientific research reference only and has not been peer-reviewed. It should not be used as the basis for public health decisions. All model results are based on theoretical assumptions and do not represent actual epidemic predictions.