main/in development

Every repo
breaks differently

Riffle ranks your pull request queue by risk, trained on your repository’s own revert history.

contractsintakescorerexplainerapp

Planned training corpus: 50 public repositories, then yours

Services

Four services

Each boundary is a place where the failure mode changes. Intake must never be slow, scoring must never be lost, and explanation is allowed to fail.

intake

The webhook front door

Verify the signature, deduplicate by delivery ID, publish, return 200 — inside the ten seconds GitHub allows. Nothing else happens here.

Explore intake
10shard budget
POST /webhook
X-GitHub-Delivery: 8f2c…a91
X-Hub-Signature-256: sha256=…
 
✓ signature verified
✓ duplicate no
→ published pr_event
200 in 41ms
scorer

Features, model, result

Consumes the event, extracts features, runs the tenant's own ranking model, asks the explainer for a sentence, writes the result.

Explore scorer
1 modelper tenant
$ riffle score --delivery 8f2c…a91
 
features 100–300 planned
model tenant:4192 v7
fallback global base
 
risk_score 0.81
rank_band review_first
model_version v7 (pinned)
explainer

Allowed to fail

LLM inference behind an API. Warm GPU, cached, rate-limited, and a hard timeout to a deterministic template. Never on the correctness path.

Explore explainer
nullvalid answer
POST /explain 900ms
→ cache miss
→ inference timeout
→ fallback template
 
explanation null
✓ rank still valid
app

The only human surface

The GitHub App and the dashboard. It reorders the queue, and it never merges a pull request or removes one from review.

Explore app
0auto-merges
queue riffle/api (14 open)
 
#2841 review_first 0.81
#2838 senior_rec. 0.64
#2844 standard 0.22
#2839 standard 0.19

Planned on fifty public repositories

Pull requests in the corpus

measured

~1,899,996

pull requests measured across 48 repositories, open and closed

By domain

measured
  • Systems & dev tools23%
  • Web frameworks14%
  • Data systems22%
  • Cloud & infra14%
  • ML & end-user27%

40 / 10

decision

repositories for training / unseen-repo validation

100–300

estimate

feature columns per pull request

0.65+

target

ROC-AUC target on repositories never trained on Kamei et al., 2016 ↗

Our claim is narrower than a verdict: the next incident follows the seams your repository has already broken along, and its own history is where they are written.

Learn more

Research this
builds on

Faros AIAcceleration Whiplash, 2026LinearB2026 BenchmarksGitHubAgent PRs are everywhereApacheJITMSR 2022MSR 2020On the Shoulders of Giants

Invariants

Breaking one is a bug even if the tests pass

  1. 01Idempotent scoringThe same delivery ID must never produce two scores.
  2. 02Version pinningA request never mixes a new model with an old feature extractor.
  3. 03Tenant fairnessOne monorepo pushing 500 PRs an hour must not starve a small team.
  4. 04Training backpressureBounded job pool. Excess requests queue or defer, never run.
  5. 05Fail open, not closedIf scoring fails the PR still appears, unranked and flagged.
  6. 06Explanation is optionalThe explainer's failure degrades quality only.

Contracts

Two shapes cross every boundary

Three languages read these schemas and nothing else defines the wire format. Any change to one is a breaking change.

intake → scorer

PrEvent

Published by intake the moment a webhook verifies — one per GitHub delivery. Its delivery_id is the idempotency key for the entire pipeline.

7 fields, all requireddelivery_id dedupes
{
"delivery_id": "X-GitHub-Delivery header",
"tenant_id": "string, installation id",
"repo": "string, owner/name",
"pr_number": 0,
"action": "opened | synchronize | reopened",
"head_sha": "string",
"received_at": "RFC3339 timestamp"
}
scorer → app

ScoreResult

Written by scorer, read by the app. The model version is pinned per request, and the feature vector travels with the score so every rank can be audited.

explanation is nullablemodel_version pinned
{
"delivery_id": "string",
"tenant_id": "string",
"pr_number": 0,
"risk_score": 0.0,
"rank_band": "review_first | standard | senior_recommended",
"model_version": "string, pinned for this request",
"features": { "...": "vector used, for audit" },
"explanation": "string | null (explainer timed out)",
"scored_at": "RFC3339 timestamp"
}

Reference & internals

All docs
ts
docs .filter(d => d.settled === true) .filter(d => d.source === "README") .filter(d => d.sections.length === 3)