Skip to content
View joesposito8's full-sized avatar

Block or report joesposito8

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
joesposito8/README.md

Joey Esposito

AI Engineer @ LinkedIn · Independent AI Safety Evaluation Engineer
San Francisco, USA


Currently an independent AI Safety Evals Engineer contributing to UK AISI's Inspect eval framework and publishing results from independent evaluations.

Projects

An eval built on Inspect attempting to measure how abliteration and prefill attacks on open-weight models unlock refusal on harmful queries, and whether they are stronger when used together.

Selected contributions

inspect_ai

  • #3709 - vLLM chat-template controls for base-model evals
  • #3969 - pass_k epoch reducer (τ-bench pass^k consistency metric)
  • #4380 - Add instance parameter and path canonicalization to memory
  • #4035 - Krippendorff's α metric for multi-judge agreement

inspect_evals

  • #1429 - fix CodeIPI exfiltration scorer to check tool-result messages
  • #1501 - cyberseceval_4: tolerate fenced / prose-wrapped judge JSON
  • #1503 - fix mean_of on_missing="skip" to also skip None-valued samples

inspect_scout

  • #455 - resolve ModelEvent input refs from the events_data pool schema

Contact

Hit me up! Open to collaboration on evals, tooling, or if you just want to contact me :)

LinkedIn · Email

Popular repositories Loading

  1. inspect_ai inspect_ai Public

    Forked from UKGovernmentBEIS/inspect_ai

    Inspect: A framework for large language model evaluations

    Python

  2. inspect_evals inspect_evals Public

    Forked from UKGovernmentBEIS/inspect_evals

    Collection of evals for Inspect AI

    Python

  3. joesposito8 joesposito8 Public

    Profile README

  4. TarantuBench TarantuBench Public

    Forked from Trivulzianus/TarantuBench

    The full repo of all the labs available as part of the benchmark

    JavaScript

  5. inspect_scout inspect_scout Public

    Forked from meridianlabs-ai/inspect_scout

    In-depth analysis of AI agent transcripts.

    Python

  6. inspect_petri inspect_petri Public

    Forked from meridianlabs-ai/inspect_petri

    An alignment auditing agent capable of quickly exploring alignment hypothesis

    Python