ResearchForge — red panda mascot
Open-source · Apache 2.0 · Python 3.12+

Your AI generates ideas.ResearchForge tests which ones actually hold up.

Claude Code or Cursor is the researcher. ResearchForge is the lab protocol.

Freeze the baseline, test competing hypotheses in isolated Git worktrees, enforce metrics and constraints, re-validate the winner, and ship with traceable evidence.

Works withClaude CodeCursorPython venvDockerGit

ResearchForge

From papers to proof in 90 seconds.

From literature to hypotheses, controlled experiments and reproducible evidence.

Product demo

What it covers

  • Research and hypothesis generation from arXiv and local code
  • Baseline locking and isolated experiment worktrees
  • Traceable lineage, rejection reasons, and validation evidence
  • Shipping the winner into a clean, reproducible branch
▶Watch on YouTube

Claude and Cursor are the researchers. ResearchForge is the lab protocol.

Coding agents are good at exploring ideas, reading papers and modifying code. But the same agent shouldn't get to define the experiment, change the implementation, and decide whether its own work succeeded.

ResearchForge separates creativity from evaluation. The agent proposes. ResearchForge freezes, executes, measures, rejects and validates.

Let agents explore. Keep evaluation deterministic.

ResearchForge

Without it vs. With it.

Most ML teams run experiments the same messy way for years. ResearchForge changes that.

Without ResearchForge
  • No systematic way to survey the relevant literature
  • No record of what failed or why
  • Main branch polluted by broken experiments
  • Baseline drifts — “improvement” is guesswork
  • Good ideas stay unvalidated because testing is slow
  • No lineage: can't reproduce last month's winner
  • The same agent proposes the change and judges whether it succeeded
With ResearchForge
  • arXiv searched end-to-end — top papers ranked and stored automatically
  • Full experiment lineage — every result, every rejection, every reason
  • Git worktrees keep your checkout untouched, always
  • Frozen baseline before anything runs — reproducible from day one
  • Validate, ship, or reject hypotheses in minutes
  • One-line handoff: ship the validated winner as a clean branch + report.
  • The agent generates the hypothesis; a deterministic protocol evaluates the result

ResearchForge

A scientific method for AI experimentation.

Four steps that turn a research question into a validated, shippable result with full lineage.

01Papers → evidence

Find the evidence

Search relevant papers and connect findings to the problem in your repository.

$ researchforge research search
02Evidence → hypotheses

Form competing hypotheses

Turn evidence into testable changes with a target metric and explicit constraints.

$ researchforge hypotheses
03Hypotheses → results

Run controlled experiments

Freeze the baseline and test every hypothesis in an isolated Git worktree.

$ researchforge run
04Results → decision

Reject, validate, ship

Keep failures in the record, retest the winner, and ship a clean branch with evidence.

$ researchforge ship

A real ResearchForge run

ROGII Kaggle · $50,000 prize pool

15.2 RMSE
baseline
7 hypotheses
formed
9 experiments
run
3 rejected
with reasons
6.1 RMSE
winner 🏆

Every experiment. Every failure. Every reason why.

ResearchForge builds an experiment graph — not a list. Winners branch into the next round. Failures stay on record. Merges combine independent gains. This is a real autorun against YOLOv5.

Round 0 — baselineRound 1Round 2Round 3Round 4 — winner
◈baselinemAP 0.7395✓exp-001Δ +0.9%○exp-002rejected✓exp-003Δ +1.4%✗exp-004291ms > budget—exp-005NO CHANGE✗exp-006exit 1✓exp-007Δ +0.7%○exp-008rejected✓exp-009Δ +2.1%MERGE—exp-010NO CHANGE★exp-011+3.9% mAP 🏆
◈Baseline
✓Improved
○Rejected
✗Failed
—NO CHANGE (inherited gain)
⋯Merge (dashed edge)
3 worktrees · running in parallel · main branch: untouched

Run experiments in parallel isolation

Every experiment gets its own isolated git worktree at the baseline commit. Your checkout is never touched — no matter how many run at once.

main branch · protected
your-repo /
baseline commit abc1234 · untouched
worktree-exp-001
PASS
# hyp: Per-well normalisation
$ researchforge run exp-001
↳ applying patch to src/
↳ venv ready · installing deps
Running benchmarks/evaluate.py...
✓ artifacts/results.json written
f1 = 0.901
Δ +4.2%
⏱ 2m 14s
NORMALIZE=True
worktree-exp-002
FAIL
# hyp: N-gram feature augmentation
$ researchforge run exp-002
↳ applying patch to src/
↳ venv ready · installing deps
Running benchmarks/evaluate.py...
✗ constraint: p95_ms=432 > 200
p95 = 432ms
✗ constraint
⏱ 3m 01s
NGRAM=True
worktree-exp-003
ERROR
# hyp: External library patch
$ researchforge run exp-003
↳ applying patch to src/
↳ venv ready · installing deps
ModuleNotFoundError: broken_lib
✗ exit code 1 · experiment failed
exit 1
fatal error
⏱ 0m 04s
import broken_lib
🔒
Zero modifications
Main branch never touched
⚡
Fully parallel
All N experiments run at once
🧹
Auto cleanup
Worktrees removed after run

ResearchForge

Give your coding agent a scientific method.

Use ResearchForge from Claude Code or Cursor to turn agent-generated ideas into controlled, reproducible experiments.

Claude Code

Skills in ~/.claude/skills/

/researchforge-start

Cursor

Rules in ~/.cursor/rules/

@researchforge-start

↕ or both at once

$ researchforge all install --user

ResearchForge

Fits around your stack. Does not replace it.

Keep your repositories, benchmarks, CI, and experiment trackers. ResearchForge adds the experimentation and evidence layer between an idea and a decision.

READS FROM
arXivGitHubYour codebaseMLflow logsProW&B runsPro
RUNS THROUGH
PythonDockerGitCI/CDProGPU runnersPro
SHIPS TO
GitHubJSONGitSlackProMLflowPro

* Enterprise tier — MLflow, W&B, CI/CD, GPU runners, Slack, custom adapters

Explore enterprise features →

ResearchForge

What the research loop looks like in practice.

Two live competitions. Real results. Traceable experiments. No cherry-picking.

🏆 $50,000 Prize Pool · Kaggle

RMSE 15.2 → 6.1

ROGII Wellbore Geology Prediction

9
experiments run
3
rejected (with reasons)
30
top papers surfaced
7
hypotheses formed

The winning hypothesis originated from a paper surfaced during the literature search phase. ResearchForge linked it to a hypothesis, ran it in isolation, and produced the exact reason it beat the baseline.

Rank ~1,700 / 6,173 teamsRead the story →
🥉 Bronze Medal · ARC Prize 2026

Rank #203 / ~2,500 teams

ARC-AGI-3 — Fluid Intelligence Benchmark

4
validated variants
1.21
best score
hyp-002
winning hypothesis
active
still competing

Score measures the fraction of ARC-AGI-3 tasks solved correctly. The public leaderboard baseline at competition start was 0.08; graph-frontier exploration (hyp-002) reached 1.21 across 4 tracked experiment variants — all with full ResearchForge lineage.

Competition closes ~2 monthsRead the story →

ResearchForge

When experimentation becomes a team workflow.

The open-source CLI proves what works locally. ResearchForge for teams brings the same reproducible loop to shared infrastructure, CI/CD, governance, and organization-wide experiment history.

🔒

Local-first, VPC-ready

Core runs locally against your repo. Teams can deploy into their own controlled environment.

📋

Complete audit trail

Every change, result and decision is traceable — lineage from paper to branch.

👥

Multi-user hub

Shared experiment dashboard and approval workflow across your whole ML team.

⚡

CI/CD integration

GitHub Actions + GitLab CI — experiments triggered and validated on every PR.

📊

MLflow / W&B bridge

ResearchForge orchestrates hypotheses and validation; your tracker stays the telemetry record.

🤖

Custom model support

Use your private fine-tuned model or provider — not locked to a single LLM.

☁️

Cloud execution

Scale experiments beyond local machines with supported cloud execution backends.

🛡️

Built for auditable environments

Protected path enforcement, reproducibility artifacts, and compliance-ready experiment records.

Bring us a benchmark.

Give us one repository, one metric and one problem your ML team wants to improve. We will show you what a ResearchForge pilot looks like.

ResearchForge

ResearchForge is free forever.

Individual researchers and open-source projects always get the full CLI — no nags, no limits.

Open Source
Free forever
  • Unlimited local experiments
  • Full experiment lineage
  • Claude Code + Cursor skills
  • Git worktree isolation
  • Community support
Get Started
Most popular
Team
Contact us
  • Multi-user hub dashboard
  • CI/CD plugin (Actions / GitLab)
  • Slack / Teams notifications
  • Email support SLA
  • SSO / SAML ready
Contact Us
Enterprise
Custom contract
  • Air-gapped / VPC deployment
  • Okta, Azure AD (SSO)
  • MLflow / W&B integration
  • SOC2 audit trail
  • Dedicated support + SLA
Book a Demo

ResearchForge

Common questions.

Answers to what most teams ask before running their first experiment.

Is ResearchForge an autonomous coding agent?

No. It is an experimentation engine and protocol used alongside coding agents. ResearchForge controls baselines, runs, constraints, and validation — the agent generates ideas, ResearchForge proves which ones work.

Why not just use Claude Code or Cursor directly?

Coding agents are great at exploration: reading code, finding papers, forming hypotheses and implementing changes. The problem is letting the same probabilistic system also decide whether its experiment succeeded. ResearchForge separates those responsibilities. Claude/Cursor stays creative; ResearchForge freezes the baseline, runs the benchmark and constraints, preserves failures, compares results and validates the winner. The agent is the researcher. ResearchForge is the lab protocol.

Does it replace MLflow or Weights & Biases?

No. ResearchForge sits above your existing tracker and focuses on research hypothesis → controlled experiment → validation lineage. Your tracker stays the system of record for run telemetry.

Does my code leave my machine?

The core CLI runs locally against your repository. Literature search calls arXiv. No experiment data is sent to external servers. Enterprise deployments run entirely inside your own environment.

What happens when an experiment fails?

The failure stays in the record with the error or violated constraint reason. Your working branch checkout is never modified — each experiment runs in its own isolated Git worktree.

Can a team share experiments?

The open-source CLI keeps experiment state locally. Teams that need a shared dashboard, approval workflow, and organization-wide lineage can deploy the Enterprise Hub into their own environment.

What kinds of benchmarks does ResearchForge support?

Any benchmark that emits a results.json with a numeric metric. Python, Docker, or any executable that writes the expected output format. See the docs for the exact contract.

Apache 2.0 · Built in the open

The research engine is open source.
And it stays that way.

Run unlimited local experiments. Apache 2.0 gives teams permission to use, modify, and build on ResearchForge — including commercially.