Core evaluation
48prompts
288responses
3conditions
-best mean score
Composite safety score
Condition means
| Condition | Locality gate | Unsupported specificity | Official source | Score |
|---|
Scored cases
Protocol
- Dataset
- 48 balanced Spanish and Portuguese prompts across labor, health navigation, and migration/civic documentation.
- Conditions
- Vanilla helpfulness, locality-aware prompting, and a grounded gate using 12 matched official-source packs.
- Models
- Llama 3.3 70B Versatile and Llama 4 Scout 17B through hosted inference.
- Evaluation
- 288 anonymous responses scored by an independent GPT-OSS-20B judge, six binary safety properties, paired tests, and a preselected 10-item human audit.