
Your AI generates ideas.ResearchForge tests which ones actually hold up.
Claude Code or Cursor is the researcher. ResearchForge is the lab protocol.
Freeze the baseline, test competing hypotheses in isolated Git worktrees, enforce metrics and constraints, re-validate the winner, and ship with traceable evidence.
ResearchForge
From papers to proof in 90 seconds.
From literature to hypotheses, controlled experiments and reproducible evidence.
What it covers
- Research and hypothesis generation from arXiv and local code
- Baseline locking and isolated experiment worktrees
- Traceable lineage, rejection reasons, and validation evidence
- Shipping the winner into a clean, reproducible branch
Claude and Cursor are the researchers. ResearchForge is the lab protocol.
Coding agents are good at exploring ideas, reading papers and modifying code. But the same agent shouldn't get to define the experiment, change the implementation, and decide whether its own work succeeded.
ResearchForge separates creativity from evaluation. The agent proposes. ResearchForge freezes, executes, measures, rejects and validates.
Let agents explore. Keep evaluation deterministic.
ResearchForge
Without it vs. With it.
Most ML teams run experiments the same messy way for years. ResearchForge changes that.
- No systematic way to survey the relevant literature
- No record of what failed or why
- Main branch polluted by broken experiments
- Baseline drifts — “improvement” is guesswork
- Good ideas stay unvalidated because testing is slow
- No lineage: can't reproduce last month's winner
- The same agent proposes the change and judges whether it succeeded
- arXiv searched end-to-end — top papers ranked and stored automatically
- Full experiment lineage — every result, every rejection, every reason
- Git worktrees keep your checkout untouched, always
- Frozen baseline before anything runs — reproducible from day one
- Validate, ship, or reject hypotheses in minutes
- One-line handoff: ship the validated winner as a clean branch + report.
- The agent generates the hypothesis; a deterministic protocol evaluates the result
ResearchForge
A scientific method for AI experimentation.
Four steps that turn a research question into a validated, shippable result with full lineage.
Find the evidence
Search relevant papers and connect findings to the problem in your repository.
$ researchforge research searchForm competing hypotheses
Turn evidence into testable changes with a target metric and explicit constraints.
$ researchforge hypothesesRun controlled experiments
Freeze the baseline and test every hypothesis in an isolated Git worktree.
$ researchforge runReject, validate, ship
Keep failures in the record, retest the winner, and ship a clean branch with evidence.
$ researchforge shipA real ResearchForge run
ROGII Kaggle · $50,000 prize pool
Every experiment. Every failure. Every reason why.
ResearchForge builds an experiment graph — not a list. Winners branch into the next round. Failures stay on record. Merges combine independent gains. This is a real autorun against YOLOv5.
Run experiments in parallel isolation
Every experiment gets its own isolated git worktree at the baseline commit. Your checkout is never touched — no matter how many run at once.
ResearchForge
Give your coding agent a scientific method.
Use ResearchForge from Claude Code or Cursor to turn agent-generated ideas into controlled, reproducible experiments.
Claude Code
Skills in ~/.claude/skills/
/researchforge-startCursor
Rules in ~/.cursor/rules/
@researchforge-start↕ or both at once
$ researchforge all install --userResearchForge
Fits around your stack. Does not replace it.
Keep your repositories, benchmarks, CI, and experiment trackers. ResearchForge adds the experimentation and evidence layer between an idea and a decision.
* Enterprise tier — MLflow, W&B, CI/CD, GPU runners, Slack, custom adapters
Explore enterprise features →ResearchForge
What the research loop looks like in practice.
Two live competitions. Real results. Traceable experiments. No cherry-picking.
RMSE 15.2 → 6.1
ROGII Wellbore Geology Prediction
The winning hypothesis originated from a paper surfaced during the literature search phase. ResearchForge linked it to a hypothesis, ran it in isolation, and produced the exact reason it beat the baseline.
Rank #203 / ~2,500 teams
ARC-AGI-3 — Fluid Intelligence Benchmark
Score measures the fraction of ARC-AGI-3 tasks solved correctly. The public leaderboard baseline at competition start was 0.08; graph-frontier exploration (hyp-002) reached 1.21 across 4 tracked experiment variants — all with full ResearchForge lineage.
ResearchForge
When experimentation becomes a team workflow.
The open-source CLI proves what works locally. ResearchForge for teams brings the same reproducible loop to shared infrastructure, CI/CD, governance, and organization-wide experiment history.
Local-first, VPC-ready
Core runs locally against your repo. Teams can deploy into their own controlled environment.
Complete audit trail
Every change, result and decision is traceable — lineage from paper to branch.
Multi-user hub
Shared experiment dashboard and approval workflow across your whole ML team.
CI/CD integration
GitHub Actions + GitLab CI — experiments triggered and validated on every PR.
MLflow / W&B bridge
ResearchForge orchestrates hypotheses and validation; your tracker stays the telemetry record.
Custom model support
Use your private fine-tuned model or provider — not locked to a single LLM.
Cloud execution
Scale experiments beyond local machines with supported cloud execution backends.
Built for auditable environments
Protected path enforcement, reproducibility artifacts, and compliance-ready experiment records.
Bring us a benchmark.
Give us one repository, one metric and one problem your ML team wants to improve. We will show you what a ResearchForge pilot looks like.
ResearchForge
ResearchForge is free forever.
Individual researchers and open-source projects always get the full CLI — no nags, no limits.
- Unlimited local experiments
- Full experiment lineage
- Claude Code + Cursor skills
- Git worktree isolation
- Community support
- Multi-user hub dashboard
- CI/CD plugin (Actions / GitLab)
- Slack / Teams notifications
- Email support SLA
- SSO / SAML ready
- Air-gapped / VPC deployment
- Okta, Azure AD (SSO)
- MLflow / W&B integration
- SOC2 audit trail
- Dedicated support + SLA
ResearchForge
Common questions.
Answers to what most teams ask before running their first experiment.
Is ResearchForge an autonomous coding agent?
No. It is an experimentation engine and protocol used alongside coding agents. ResearchForge controls baselines, runs, constraints, and validation — the agent generates ideas, ResearchForge proves which ones work.
Why not just use Claude Code or Cursor directly?
Coding agents are great at exploration: reading code, finding papers, forming hypotheses and implementing changes. The problem is letting the same probabilistic system also decide whether its experiment succeeded. ResearchForge separates those responsibilities. Claude/Cursor stays creative; ResearchForge freezes the baseline, runs the benchmark and constraints, preserves failures, compares results and validates the winner. The agent is the researcher. ResearchForge is the lab protocol.
Does it replace MLflow or Weights & Biases?
No. ResearchForge sits above your existing tracker and focuses on research hypothesis → controlled experiment → validation lineage. Your tracker stays the system of record for run telemetry.
Does my code leave my machine?
The core CLI runs locally against your repository. Literature search calls arXiv. No experiment data is sent to external servers. Enterprise deployments run entirely inside your own environment.
What happens when an experiment fails?
The failure stays in the record with the error or violated constraint reason. Your working branch checkout is never modified — each experiment runs in its own isolated Git worktree.
Can a team share experiments?
The open-source CLI keeps experiment state locally. Teams that need a shared dashboard, approval workflow, and organization-wide lineage can deploy the Enterprise Hub into their own environment.
What kinds of benchmarks does ResearchForge support?
Any benchmark that emits a results.json with a numeric metric. Python, Docker, or any executable that writes the expected output format. See the docs for the exact contract.
The research engine is open source.
And it stays that way.
Run unlimited local experiments. Apache 2.0 gives teams permission to use, modify, and build on ResearchForge — including commercially.