- Fresno, CA
-
04:53
(UTC -07:00) - in/redswimmer
Pinned Loading
-
wikipedia-qa-agent
wikipedia-qa-agent PublicEvaluating an agent for answer quality and safety that uses Claude and Wikipedia to answer questions.
Python 1
-
turn-level-rewards
turn-level-rewards PublicReproducing a turn-level vs. outcome reward ablation for RL fine-tuning of multi-turn search agents (GRPO/PPO), based on arXiv:2505.11821
Python 1
-
ml-internal-llms-instruction-following
ml-internal-llms-instruction-following PublicForked from apple/ml-internal-llms-instruction-following
Do LLMs know when to say no? Extending Apple's ICLR 2025 instruction-following probes to agentic tool calling.
Python 1
-
trace-inversion
trace-inversion PublicRecreation of How to Steal Reasoning Without Reasoning Traces, based on arXiv:2603.07267
Python 1
If the problem persists, check the GitHub status page or contact support.