We catch reward exploits
before models learn them.
any% is a small research team focused on reward hacking in RLVR environments.
We study failures in real training runs and build tools to find, understand, and prevent them.
- max_steps10,000
- rollouts_per_example16
- seq_len4096
- batch_size512
- max_tokens2048
- learning_rate1e-6
Current focus
Auditing open-source RLVR envs from Prime Intellect's Hub.One env at a time, we try to find shortcuts that earn reward without solving the task, building the tooling we need along the way.
Talk to us
Found something strange?
If you have an RLVR environment or training run that seems to score well for the wrong reasons, send it our way. We'll work out what the model found, why it earned reward, and what needs to change.
Want to work on this with us?
We're looking for people who want to become unusually good at understanding and preventing reward hacking. If you're interested, tell us what you've worked on and what you'd like to investigate.