Decomposing and measuring evaluation awareness in existing benchmarks and our proposed EvalAwareBench.
-
Updated
Jun 1, 2026 - Python
Decomposing and measuring evaluation awareness in existing benchmarks and our proposed EvalAwareBench.
Repository for the arXiv 2026 prepring "Models That Know How Evaluations Are Designed Score Safer"
Synthetic-document fine-tuning on Qwen2.5-7B: a controlled study of whether SDF installs sandbagging, finding a layered recognition/generation/behavior dissociation.
Agentic evaluation of misalignment, scheming and evaluation awareness in LLM agents. Scenario text encoded at rest to resist training-data contamination.
Certified reachability of NLA-defined concepts: proving when a model can(not) recognize it is being tested. Mashes Anthropic's Natural Language Autoencoder with Raghunathan certified defenses, on GPT-2.
Testing which evaluation-context cues change Qwen3-32B's task-completion behavior.
[WIP] Code and datasets for the thesis "Sensitivity Analysis of Evaluation Awareness in Large Language Models" · Licenciatura en Ciencia de Datos, Universidad de Buenos Aires (UBA).
To associate your repository with the evaluation-awareness topic, visit your repo's landing page and select "manage topics."