Skip to content

docs: add RLHF reward hacking evaluation lesson - #157

Open
AKilalours wants to merge 1 commit into
anthropics:masterfrom
AKilalours:docs/rlhf-reward-hacking-eval
Open

docs: add RLHF reward hacking evaluation lesson#157
AKilalours wants to merge 1 commit into
anthropics:masterfrom
AKilalours:docs/rlhf-reward-hacking-eval

Conversation

@AKilalours

Copy link
Copy Markdown

Summary

This PR adds a lightweight RLHF reward-hacking evaluation lesson to the prompt evaluations materials.

Why

Preference optimization can improve model behavior, but it can also create failure modes where models optimize for surface-level reward signals rather than the underlying objective. This lesson introduces evaluation patterns for sycophancy, unsafe helpfulness, overconfidence, and unsupported confidence.

Changes

  • Added an RLHF reward-hacking evaluation lesson
  • Added preference-pair examples for alignment-sensitive behavior
  • Added a scoring rubric for reward-hacking failure modes
  • Added a suggested evaluation workflow for regression testing
  • Linked the lesson from the prompt evaluations README

Testing

  • Documentation-only change
  • Checked Markdown formatting locally
  • Commit signed and verified

@AKilalours

Copy link
Copy Markdown
Author

Hi maintainers,

I opened this as a lightweight prompt-evaluations lesson focused on RLHF reward-hacking failure modes, including sycophancy, unsafe helpfulness, overconfidence, and unsupported confidence.

Happy to revise the wording, placement, or scope if you prefer a different structure for the course materials.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant