experiment-audit

by wanshuiyin 15k MIT Updated Aug 18, 2026
experiment-audit skill by wanshuiyin
experiment-audit — Data skill by wanshuiyin

experiment-audit is an open-source data skill for Claude Code and compatible agents, published by wanshuiyin. Its author describes it as: “Audit experiment integrity before claiming results. Uses cross-model review (external reviewer backend) to check for fake ground truth, score normalization fraud, phantom results, and insufficient scope. Use when user…”. The project has 15k stars on GitHub and is available under the MIT license. Add it to your setup with `git clone https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep ~/.claude/skills/experiment-audit`.

What experiment-audit does

> 🔒 **Do not wrap this skill in `/loop`, `/schedule`, or `CronCreate`.** It is > verdict-bearing — it judges experiment integrity. Re-running that verdict on a > timer adds no new signal, and a loop that accepts its own output to decide > when to stop crosses into self-acquittal (`acceptance-gate.md`). Schedule the > *external wait that precedes it* — experiments done → then audit **once**. See > `shared-references/external-cadence.md`.

Installation

Add experiment-audit to your agent with:

git clone https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep ~/.claude/skills/experiment-audit

Always review a skill's source before installing it. This command comes from the skill's public repository; the linked repo is the source of truth for exact setup steps.

What's inside

The SKILL.md for experiment-audit is organised into these sections:

  • Why This Exists
  • Core Principle
  • Constants
  • Reviewer Calling Convention
  • Workflow
  • Step 1: Collect Artifacts (Executor — Claude)
  • Step 2: Send to Reviewer
  • Step 3: Parse and Write Report (Executor — Claude)
  • Step 4: Print Summary
  • Integration with Other Skills
  • Automatic in /research-pipeline (advisory, never blocks)
  • Read by /result-to-claim (if exists)

When to use it

Reach for experiment-audit when you want data help from your agent without writing the same instructions every session. Load the skill and the agent picks it up automatically for relevant tasks.

Strengths

  • Clear MIT license — safe to read and adapt
  • Ships in wanshuiyin/Auto-claude-code-research-in-sleep, an established project with 14,858 GitHub stars
  • Actively maintained (recent commits)

Topics

ai-researchai-toolsarisautonomous-agentclaudeclaude-codeclaude-code-skillscodexdeep-learninggptidea-generationllmmachine-learningmcpmcp-serverml-researchopenaipaper-reviewpaper-writingresearch-automation

Frequently asked questions

What does experiment-audit do?
Audit experiment integrity before claiming results. Uses cross-model review (external reviewer backend) to check for fake ground truth, score normalization fraud, phantom results, and insufficient scope. Use when user says "审计实验", "check experiment integrity", "audit results", "实验诚实度", or after experiments complete before writing claims.
How do I install experiment-audit?
Run git clone https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep ~/.claude/skills/experiment-audit in your agent, then reload your skills. Review the source at https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep before installing.
Is experiment-audit free to use?
Yes. experiment-audit is free and open source under the MIT license, so you can read, run, and adapt it within that license's terms.
Where does experiment-audit come from?
experiment-audit ships inside wanshuiyin/Auto-claude-code-research-in-sleep, a repository that contains 41 catalogued skills in total. The repository's 14,858 GitHub stars apply to that whole collection, not to this skill on its own.

Related skills

More Data →

Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.

15k wanshuiyin MIT

Quick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.

15k wanshuiyin MIT

Analyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says "analyze results", "compare", or needs to interpret experimental data.

15k wanshuiyin MIT

Search, download, and summarize academic papers from arXiv. Use when user says "search arxiv", "download paper", "fetch arxiv", "arxiv search", "get paper pdf", or wants to find and save papers from arXiv to the local paper library.

15k wanshuiyin MIT

Autonomously improve a generated paper via GPT-5.6-Sol xhigh review → implement fixes → recompile, for 2 rounds. Use when user says "改论文", "improve paper", "论文润色循环", "auto improve", or wants to iteratively polish a generated paper.

15k wanshuiyin MIT