Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: 037834f4
Paper published March 1, 2026
Entry analyzed March 8, 2026
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

Thai Cultural Alignment Bypass

A vulnerability in the safety alignment of Large Language Models (LLMs) allows attackers to bypass safety guardrails by using malicious prompts contextualized in the Thai language and culture. Evaluated models exhibit…

BibTeX citation

Paper-evaluated models(15)

Qwen 2.5 7B Instruct, Qwen 2.5 72B Instruct, Llama 3.1 8B Instruct +12 more
  • Qwen 2.5 7B Instruct
  • Qwen 2.5 72B Instruct
  • Llama 3.1 8B Instruct
  • Llama 3.2 1B Instruct
  • Llama 3.3 70B Instruct
  • Gemma 3 4B IT
  • Gemma 3 12B IT
  • SeaLLMs v3 1.5B
  • SeaLLMs v3 7B
  • Llama SEA-LION v3 8B
  • Llama SEA-LION v3 70B
  • Typhoon 2
  • Typhoon 2.1
  • OpenThaiGPT 1.5 7B
  • OpenThaiGPT 1.5 72B

Description

A vulnerability in the safety alignment of Large Language Models (LLMs) allows attackers to bypass safety guardrails by using malicious prompts contextualized in the Thai language and culture. Evaluated models exhibit a significantly higher Attack Success Rate (ASR) against Thai-specific, culturally contextualized attacks compared to general translated attacks. By exploiting local cultural nuances, regional slang, and Thai socio-cultural contexts, attackers can easily circumvent standard safety filters to elicit harmful responses, particularly in the domain of Thai socio-cultural harms where model performance is notably weaker.

Examples

See the ThaiSafetyBench dataset on HuggingFace and GitHub for 1,954 explicit examples of culturally contextualized jailbreaks, including specific prompts utilizing Thai slang, localized fake news narratives, and deliberate violations of Thai social etiquette.

Impact

Successful exploitation allows attackers to bypass standard LLM safety mechanisms to generate prohibited content. This includes generating culturally-specific hate speech, unfair discrimination, sensitive organizational information leakage, and the dissemination of localized misinformation (e.g., narratives reflecting prevalent fake news in Thai society).

Affected Systems

Various open-source multilingual and regionally-tuned LLMs, including but not limited to:

  • Qwen2.5 (7B, 72B Instruct)
  • Llama-3.1, 3.2, and 3.3 variants (Instruct)
  • Gemma-3 (4B, 12B IT)
  • SeaLLMs-v3 (1.5B, 7B)
  • Llama-SEA-LION-v3 (8B, 70B)
  • Typhoon2 and Typhoon2.1 variants (1B, 3B, 4B, 8B, 12B, 70B)
  • OpenThaiGPT1.5 (7B, 72B)

Mitigation Steps

  • Implement culturally tailored safety tuning that specifically addresses local cultural and contextual nuances, linguistic patterns, and region-specific slang.
  • Carefully curate Continual Pretraining (CPT) data to integrate region-specific safety datasets rather than relying solely on translated English benchmarks.
  • Enhance adversarial filtering during the training and tuning processes to align with region-specific safety requirements and socio-cultural norms.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
Black-box model, service, or application access.
Related deployment categories
No related deployment category is classified.
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
Various open-source multilingual and regionally-tuned LLMs, including but not limited to: Qwen2.5 (7B, 72B Instruct) Llama-3.1, 3.2, and 3.3 variants (Instruct) Gemma-3 (4B, 12B IT) SeaLLMs-v3 (1.5B, 7B)…

Research Paper

ThaiSafetyBench: Assessing Language Model Safety in Thai Cultural Contexts

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2603.04992