Starred repositories
Tooling that makes it easy for agents to use Runpod
Clean public RSS endpoints for blogs whose own feeds are missing, stale, broken or noisy
Curated Codex plugin marketplace for Spark — browse and install Spark CLI skills for email, calendar, contacts, and meetings
Bend 2: a fast language that blocks AI mistakes via proof. Install: curl -fsSL https://bend-lang.com/install.sh | sh
A coding-agent skill for multi-phase security audits with independently verified, machine-readable findings
Secure, fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in multi-language ruleset (N…
Python library and shell utilities to monitor filesystem events.
Open source evals for physical AI. Run any LLM/VLA on any arm/humanoid against any real/sim benchmark.
Lean certificates accompanying Navier-Stokes and Euler results
A Go implementation of the Model Context Protocol (MCP), enabling seamless integration between LLM applications and external data sources and tools.
A modern and intuitive terminal-based text editor
This is MCP server for Claude that gives it terminal control, file system search and diff file editing capabilities
ASD-STE100 Simplified Technical English rules, repurposed as a Claude Code skill for rewriting ambiguous agent-facing English.
Close-reading workbench for LLM experiment outputs — JSONL/CSV/inspect-ai .eval viewer with marks, judges, SQL, and an embedded Claude session
Manages parallel Claude Code sessions, so you don't have to.
Easily and securely send things from one computer to another 🐊 📦
OpenAI's Codex Security CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities. npm: https://www.npmjs.com/package/@openai/codex-security
Official Inspect Implementation for "ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases"
Simple, unified interface to multiple Generative AI providers
A library for making RepE control vectors
Catch and fix the code-quality issues AI coding agents leave behind - dead code, unsafe casts, swallowed errors, duplication, security risks, and more. 50+ deterministic rules across 10 language ta…
Show usage stats for OpenAI Codex and Claude Code, without having to login.
EdgeBench: Unveiling scaling laws of learning from real-world environments
Companion code for the global workspace interpretability paper