DynaGuard: A dynamic guardian model with user-defined policies
A suite of dynamic guardian models offering novel flexibility by evaluating text based on user-defined policies.
Guardian models play a crucial role in ensuring the safety and ethical behavior of user-facing AI applications by enforcing guardrails and detecting harmful content. While standard guardian models are limited to predefined, static harm categories, we introduce DynaGuard, a suite of dynamic guardian models offering novel flexibility by evaluating text based on user-defined policies, and DynaBench, a dataset for training and evaluating dynamic guardian models. Our models provide both rapid detection of policy violations and a chain-of-thought reasoning option that articulate and justify model outputs. Critically, DynaGuard not only surpasses static models in detection accuracy on traditional safety categories, but is competitive with frontier reasoning models on free-form policy violations, all in a fraction of the time. This breakthrough makes DynaGuard a critical tool for language model guardrails.
Latest publications
Alignment-weighted DPO
A DPO that targets the most problematic parts of an output by assigning different preference weights.
ICLRYour model diversity determines reasoning strategy
A framework decomposing reasoning uncertainty and deriving conditions where depth refinement outperforms parallel sampling. (ICLR)
ICLRMR3: Multilingual rubric-agnostic reward reasoning models
A multilingual, rubric-agnostic reward reasoning model achieving the broadest language coverage in reward modeling to date.
ICLR