Skip to content
@safety-research

Safety Research

Popular repositories Loading

  1. bloom bloom Public

    bloom - evaluate any behavior immediately  🌸🌱

    Python 1.4k 171

  2. persona_vectors persona_vectors Public

    Persona Vectors: Monitoring and Controlling Character Traits in Language Models

    Python 465 114

  3. automated-w2s-research automated-w2s-research Public

    Python 311 48

  4. SCONE-bench SCONE-bench Public

    187 31

  5. assistant-axis assistant-axis Public

    The Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's behavior is. Models can drift away from the Assistant during conversations—sometimes toward bizarr…

    Jupyter Notebook 174 48

  6. safety-tooling safety-tooling Public

    Inference API for many LLMs and other useful tools for empirical research

    Python 137 49

Repositories

Showing 10 of 56 repositories
  • failure-disclosure Public

    Code and data for the paper "When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning"

    safety-research/failure-disclosure's past year of commit activity
    0 Apache-2.0 0 0 0 Updated Sep 27, 2026
  • safety-research/red-teaming-auto-mode's past year of commit activity
    Python 4 MIT 0 0 0 Updated Sep 25, 2026
  • safety-research/misalignment-in-ai-organizations's past year of commit activity
    Python 0 MIT 0 0 0 Updated Sep 13, 2026
  • self-modeling-train Public

    The training codebase to improve LLM self-modeling capability.

    safety-research/self-modeling-train's past year of commit activity
    Python 1 MIT 0 0 0 Updated Aug 25, 2026
  • self-modeling-eval Public

    An eval benchmark suite that tests how well a LLM can predict its own behavior.

    safety-research/self-modeling-eval's past year of commit activity
    Python 1 MIT 0 0 0 Updated Aug 25, 2026
  • safety-research/auditing-agents's past year of commit activity
    Python 33 14 1 2 Updated Jul 1, 2026
  • misalignment-indicators Public

    Source code for the paper: Probing the Misaligned Thinking Process of Language Models

    safety-research/misalignment-indicators's past year of commit activity
    Python 11 1 0 0 Updated Jun 22, 2026
  • safety-tooling Public

    Inference API for many LLMs and other useful tools for empirical research

    safety-research/safety-tooling's past year of commit activity
    Python 137 MIT 49 13 20 Updated May 29, 2026
  • SCONE-bench Public
    safety-research/SCONE-bench's past year of commit activity
    187 MIT 31 5 0 Updated May 22, 2026
  • sleight-bench Public

    Benchmark dataset for evaluating trusted monitors on AI agent transcripts

    safety-research/sleight-bench's past year of commit activity
    Python 16 MIT 2 1 0 Updated May 16, 2026

Most used topics

Loading…