SaFoLab : Security and Safe Foundation Model Systems
Pinned Loading
Repositories
- AGrail4Agent Public
[ACL 2025] The official code for "AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection".
- Self_Adapt_Harness Public
Official Repo for "HarnessRL: Self Adaptive Harness Learning for Test Time Discovery"
- ROM Public
The official implementation of our paper "ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention"
- ReasoningBomb Public
[CCS 2026] The official implementation of our CCS 2026 paper "ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models"
- DRIFT Public
[NeurIPS 2025] The official implementation of the paper "DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents".
- Latent_Policy_Guard Public
Latent Policy Guard (LPG) — a guardrail model that performs semantic latent deliberation over dynamic safety policies. LPG compresses intent and risk reasoning into latent tokens and emits a compact policy-indexed verdict.
- MaskForge Public
- SafeVL Public
Official Repo for Paper: SafeVL: Driving Safety Evaluation via Meticulous Reasoning in Vision Language Models
- PW-OPSD Public
The official implementation of our preprint paper "When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning"
- Safety-Midtrain Public
Code for Safety Mid-Training: Internalizing LLM Safety as a Foundational Capability
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…