Riffle ranks your pull request queue by risk, trained on your repository’s own revert history.
Planned training corpus: 50 public repositories, then yours
Services
Each boundary is a place where the failure mode changes. Intake must never be slow, scoring must never be lost, and explanation is allowed to fail.
Verify the signature, deduplicate by delivery ID, publish, return 200 — inside the ten seconds GitHub allows. Nothing else happens here.
Explore intakePOST /webhookX-GitHub-Delivery: 8f2c…a91X-Hub-Signature-256: sha256=…✓ signature verified✓ duplicate no→ published pr_event200 in 41ms
Consumes the event, extracts features, runs the tenant's own ranking model, asks the explainer for a sentence, writes the result.
Explore scorer$ riffle score --delivery 8f2c…a91features 100–300 plannedmodel tenant:4192 v7fallback global baserisk_score 0.81rank_band review_firstmodel_version v7 (pinned)
LLM inference behind an API. Warm GPU, cached, rate-limited, and a hard timeout to a deterministic template. Never on the correctness path.
Explore explainerPOST /explain 900ms→ cache miss→ inference timeout→ fallback templateexplanation null✓ rank still valid
The GitHub App and the dashboard. It reorders the queue, and it never merges a pull request or removes one from review.
Explore appqueue riffle/api (14 open)#2841 review_first 0.81#2838 senior_rec. 0.64#2844 standard 0.22#2839 standard 0.19
Pull requests in the corpus
measured~1,899,996
pull requests measured across 48 repositories, open and closed
By domain
measured40 / 10
decisionrepositories for training / unseen-repo validation
100–300
estimatefeature columns per pull request
0.65+
targetROC-AUC target on repositories never trained on Kamei et al., 2016 ↗
Research this
builds on
Invariants
Contracts
Three languages read these schemas and nothing else defines the wire format. Any change to one is a breaking change.
Published by intake the moment a webhook verifies — one per GitHub delivery. Its delivery_id is the idempotency key for the entire pipeline.
{"delivery_id": "X-GitHub-Delivery header","tenant_id": "string, installation id","repo": "string, owner/name","pr_number": 0,"action": "opened | synchronize | reopened","head_sha": "string","received_at": "RFC3339 timestamp"}
Written by scorer, read by the app. The model version is pinned per request, and the feature vector travels with the score so every rank can be audited.
{"delivery_id": "string","tenant_id": "string","pr_number": 0,"risk_score": 0.0,"rank_band": "review_first | standard | senior_recommended","model_version": "string, pinned for this request","features": { "...": "vector used, for audit" },"explanation": "string | null (explainer timed out)","scored_at": "RFC3339 timestamp"}