Skip to content
View redswimmer's full-sized avatar
💭
Taming the AIs
💭
Taming the AIs

Block or report redswimmer

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. turn-level-rewards turn-level-rewards Public

    Reproducing a turn-level vs. outcome reward ablation for RL fine-tuning of multi-turn search agents (GRPO/PPO), based on arXiv:2505.11821

    Python 1

  2. ml-internal-llms-instruction-following ml-internal-llms-instruction-following Public

    Forked from apple/ml-internal-llms-instruction-following

    Do LLMs know when to say no? Extending Apple's ICLR 2025 instruction-following probes to agentic tool calling.

    Python 1

  3. wikipedia-qa-agent wikipedia-qa-agent Public

    Evaluating an agent for answer quality and safety that uses Claude and Wikipedia to answer questions.

    Python 1