Dif Cloud

Experiments live in your repo. Cloud gives the team the live read.

Dif Cloud reads your repo and turns every .md experiment into a live view: exposures, lift, confidence intervals, and a ready-to-conclude flag on one page. Every change it proposes comes back as a pull request, so git stays the source of truth.

Dif Cloud dashboard showing live experiments across surfaces

Drill into one experiment. Ship when it's earned.

Open any experiment for lift over time with a 95% confidence interval, control and variant compared on exposures and rate, and the hypothesis, success criteria, and audience kept next to the numbers.

When a variant clears the bar, Cloud writes the verdict into the file's ## Decision block and opens a pull request: ship Variant A, activation up +4.8% at 96% confidence over 48,210 exposures, no drop in day-1 funding. You merge it, or you keep it running.

Experiment detail with ship-the-winner verdict and lift-over-time chart

Every experiment, grouped by where it runs.

Cloud mirrors your surfaces/ folder: onboarding, home, transfers, each with its live, queued, and concluded counts.

Every surface carries its known landmines (never gate the first deposit behind full KYC, −12%) and a running learning log (a 3-step progress bar lifted completion, +2.1%), so the next draft starts from what the last test found.

Experiments grouped by surface with landmines and learnings

Every decision keeps the learning that made it.

A log of concluded experiments with the outcome, the lift, and a one-line learning for each. A 65% win rate this quarter, one learning apiece.

When someone new asks why the cash-bonus upsell was held on control, the record answers — it eroded trust scores, −0.8% — even after the person who ran it has moved on.

Decision history log with win rate and outcomes

One metrics catalog the whole team shares.

Every metric the SDK fires, where it lives, how often it triggers, and which experiments depend on it — 24 defined once and used anywhere.

Guardrail metrics are flagged and watched for drift, so an activation win (+6.2%) that quietly drags the average transfer (−1.1%) gets caught before it ships.

Shared metrics catalog with guardrail tracking

Cloud drafts the next experiment from the log.

Dif reads the surface logs and the concluded history, then drafts the next test: a hypothesis, a primary metric, an audience, and the reasoning behind it.

"Prefill the transfer amount from the user's last transfer," because 34% abandon on the amount step — expected +3.4% at high confidence, drawn from three comparable prefill wins. Copy the brief and launch, or dismiss it. You always get the last word.

AI-drafted experiment suggestion with hypothesis and expected lift

Cloud proposes. You merge.

Nothing edits your experiments behind your back. The decision it drafts, the metric it adds, the experiment it scaffolds: each one arrives as a pull request against your repo, and the activity feed logs every sync.

Your experiments are still just files.

experiments/id-check.md PR #218

Conclude id-check: ship Variant A

@@ ## Decision
+ Ship Variant A. Activation up +4.8% on
+ activation_rate at 96% confidence across
+ 48,210 exposures. No drop in day-1 funding.
+ Concluded 2026-07-08, drafted by dif AI.
proposed by dif AI, ready for your review

Put the team on the same page as the repo.

Connect a repo and Cloud reads your experiments in minutes. $50 a month includes 1 million events and every seat. The CLI and SDK stay free.