The LMVD-ID is an internal research identifier, not an official CVE identifier.
Thai Cultural Alignment Bypass
A vulnerability in the safety alignment of Large Language Models (LLMs) allows attackers to bypass safety guardrails by using malicious prompts contextualized in the Thai language and culture. Evaluated models exhibit…
Paper-evaluated models(15)
Qwen 2.5 7B Instruct, Qwen 2.5 72B Instruct, Llama 3.1 8B Instruct +12 more
- Qwen 2.5 7B Instruct
- Qwen 2.5 72B Instruct
- Llama 3.1 8B Instruct
- Llama 3.2 1B Instruct
- Llama 3.3 70B Instruct
- Gemma 3 4B IT
- Gemma 3 12B IT
- SeaLLMs v3 1.5B
- SeaLLMs v3 7B
- Llama SEA-LION v3 8B
- Llama SEA-LION v3 70B
- Typhoon 2
- Typhoon 2.1
- OpenThaiGPT 1.5 7B
- OpenThaiGPT 1.5 72B
Description
A vulnerability in the safety alignment of Large Language Models (LLMs) allows attackers to bypass safety guardrails by using malicious prompts contextualized in the Thai language and culture. Evaluated models exhibit a significantly higher Attack Success Rate (ASR) against Thai-specific, culturally contextualized attacks compared to general translated attacks. By exploiting local cultural nuances, regional slang, and Thai socio-cultural contexts, attackers can easily circumvent standard safety filters to elicit harmful responses, particularly in the domain of Thai socio-cultural harms where model performance is notably weaker.
Examples
See the ThaiSafetyBench dataset on HuggingFace and GitHub for 1,954 explicit examples of culturally contextualized jailbreaks, including specific prompts utilizing Thai slang, localized fake news narratives, and deliberate violations of Thai social etiquette.
Impact
Successful exploitation allows attackers to bypass standard LLM safety mechanisms to generate prohibited content. This includes generating culturally-specific hate speech, unfair discrimination, sensitive organizational information leakage, and the dissemination of localized misinformation (e.g., narratives reflecting prevalent fake news in Thai society).
Affected Systems
Various open-source multilingual and regionally-tuned LLMs, including but not limited to:
- Qwen2.5 (7B, 72B Instruct)
- Llama-3.1, 3.2, and 3.3 variants (Instruct)
- Gemma-3 (4B, 12B IT)
- SeaLLMs-v3 (1.5B, 7B)
- Llama-SEA-LION-v3 (8B, 70B)
- Typhoon2 and Typhoon2.1 variants (1B, 3B, 4B, 8B, 12B, 70B)
- OpenThaiGPT1.5 (7B, 72B)
Mitigation Steps
- Implement culturally tailored safety tuning that specifically addresses local cultural and contextual nuances, linguistic patterns, and region-specific slang.
- Carefully curate Continual Pretraining (CPT) data to integrate region-specific safety datasets rather than relying solely on translated English benchmarks.
- Enhance adversarial filtering during the training and tuning processes to align with region-specific safety requirements and socio-cultural norms.
Research context and confidence
- Evidence and verification
- Paper-reported; independent reproduction is not documented.
- Primary research source linked.
- Severity
- Not rated by this catalog.
- Source and publication type
- arXiv · Research preprint.
- Peer-review status is not provided by this source.
- Author and publication status
- Author metadata is not stored; see the primary paper.
- Threat model and attacker access
- Black-box model, service, or application access.
- Related deployment categories
- No related deployment category is classified.
- Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
- Affected systems
- Various open-source multilingual and regionally-tuned LLMs, including but not limited to: Qwen2.5 (7B, 72B Instruct) Llama-3.1, 3.2, and 3.3 variants (Instruct) Gemma-3 (4B, 12B IT) SeaLLMs-v3 (1.5B, 7B)…
Research Paper
ThaiSafetyBench: Assessing Language Model Safety in Thai Cultural Contexts
Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.
View PaperEvidence
This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.
https://arxiv.org/abs/2603.04992Related research
- LLM Judge Framing Bias
Published January 1, 2026 · model-layer, prompt-layer, hallucination
- Selective Hate Speech Jailbreak
Published January 1, 2026 · model-layer, jailbreak, blackbox
- Linguistic Style Jailbreak
Published November 1, 2025 · prompt-layer, jailbreak, blackbox