The LMVD-ID is an internal research identifier, not an official CVE identifier.
LLM Watermark Translation Bypass
Token-level embedding-time watermarking algorithms, specifically KGW (Kirchenbauer et al.) and Exponential Sampling (EXP, Kuditipudi et al.), when implemented in Large Language Models (LLMs) for Bangla text generation…
Paper-evaluated models(1)
- Llama 3 8B
Description
Token-level embedding-time watermarking algorithms, specifically KGW (Kirchenbauer et al.) and Exponential Sampling (EXP, Kuditipudi et al.), when implemented in Large Language Models (LLMs) for Bangla text generation, are vulnerable to watermark erasure via cross-lingual round-trip translation (RTT) attacks. While these methods achieve high detection accuracy (>88%) under benign conditions, translating watermarked Bangla text to English and back to Bangla causes detection accuracy to collapse to approximately 9–13%. The vulnerability stems from the specific linguistic properties of Bangla (rich morphology, flexible word order) combined with the RTT process, which induces extensive lexical substitution and syntactic reordering. This structural disruption obliterates the token-level statistical biases required for watermark verification while preserving semantic meaning, effectively "laundering" the text.
Examples
- Setup: Deploy an instruction-tuned Bangla LLaMA-3-8B model integrated with KGW watermarking (using standard hyperparameters, e.g., green-list ratio $\gamma=0.25$).
- Generation: Generate a response to a prompt (e.g., from the Bangla-Alpaca Orca dataset).
- Result: The generated text triggers a positive detection (Z-score > 4.0).
- Attack (Text Laundering):
- Translate the generated Bangla output to English using a translation model (e.g., BanglaNMT).
- Translate the resulting English text back to Bangla.
- Verification Failure:
- Run the KGW detection algorithm on the RTT-processed text.
- Result: The Z-score drops significantly below the detection threshold, resulting in a false negative (the text is identified as human-written).
Impact
This vulnerability allows malicious actors to completely bypass authorship attribution, intellectual property protection, and AI-misuse detection systems in low-resource language contexts. It enables the undetected proliferation of AI-generated content for plagiarism, disinformation, or spam by effectively removing provenance signals without requiring model access or retraining.
Affected Systems
- Large Language Models generating Bangla text (e.g., Bangla LLaMA-3-8B).
- Implementations of KGW (Kirchenbauer et al., 2023) and Exponential Sampling (Kuditipudi et al., 2023) watermarking schemes applied to low-resource, morphologically rich languages.
Mitigation Steps
- Implement Layered Watermarking: Combine embedding-time watermarking (KGW/EXP) with a secondary, post-generation watermarking layer (such as Waterfall/Lau et al.). This creates an orthogonal statistical signal that persists even when token-level statistics are disrupted.
- Weighted Selection: During the post-generation phase, select paraphrase candidates by maximizing a weighted combination of semantic similarity and watermark strength to balance robustness with text quality.
- Orthogonal Signal Injection: Ensure the secondary layer relies on structural or distributional perturbations independent of the specific token choices made by the primary layer.
Research context and confidence
- Evidence and verification
- Paper-reported; independent reproduction is not documented.
- Primary research source linked.
- Severity
- Not rated by this catalog.
- Source and publication type
- arXiv · Research preprint.
- Peer-review status is not provided by this source.
- Author and publication status
- Author metadata is not stored; see the primary paper.
- Threat model and attacker access
- Black-box model, service, or application access.
- Related deployment categories
- No related deployment category is classified.
- Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
- Affected systems
- Large Language Models generating Bangla text (e.g., Bangla LLaMA-3-8B). Implementations of KGW (Kirchenbauer et al., 2023) and Exponential Sampling (Kuditipudi et al., 2023) watermarking schemes applied to…
Research Paper
BanglaLorica: Design and Evaluation of a Robust Watermarking Algorithm for Large Language Models in Bangla Text Generation
Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.
View PaperEvidence
This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.
https://arxiv.org/abs/2601.04534Related research
- Sophisticated Deception Induces Misbelief
Published January 1, 2026 · prompt-layer, model-layer, injection
- Back-Translation Watermark Stripping
Published November 1, 2025 · model-layer, jailbreak, blackbox
- LLM Judge Manipulation
Published March 1, 2026 · model-layer, application-layer, prompt-layer