Provenance
Co-Authored-By trailers, agent bot accounts, CLAUDE.md, .cursorrules, AGENTS.md and builder markers from Lovable, v0, Bolt and Replit. Anchored on emails and tool spellings, so a human co-author named Claude does not match.
Give Halo any link. It returns a calibrated estimate of how likely the work behind it was AI-generated, AI-assisted or human-written, with every claim pointed at a file, commit or line.
A trailer proves a tool took part in a commit, not how much it wrote. So the estimate is built in three layers, and each one is reported on its own.
Co-Authored-By trailers, agent bot accounts, CLAUDE.md, .cursorrules, AGENTS.md and builder markers from Lovable, v0, Bolt and Replit. Anchored on emails and tool spellings, so a human co-author named Claude does not match.
Commit timing, history mess, message style, account behavior, stylometry on tree-sitter ASTs for 10 languages, prose statistics and drift from the author's own pre-2021 writing. A detector without enough data never guesses.
Weighted log-odds learned by regularized logistic regression, calibrated with Platt or isotonic regression, with a 90 percent bootstrap interval and thresholds set for a low false-positive rate.
Each detector returns a score, a confidence, an applicability flag and evidence with a pointer, plus the raw features so a reader can audit the number.
Seven detectors were driven to zero by the data. Reports name them in their limitations, and a target whose only applicable detectors carry no weight is reported as insufficient_data instead of a number.
84 real GitHub repositories from 78 owners, cross-validated in 5 folds grouped by owner. Provenance was masked and footprint files removed, so these numbers measure inference alone.
| Threshold | Tier | True positives | False positives |
|---|---|---|---|
| 0.60 | mixed_or_assisted | 82.5% | 0 of 40 |
| 0.70 | none | 77.5% | 0 of 40 |
| 0.85 | mostly_ai | 60.0% | 0 of 40 |
The hosted endpoint streams each stage as a line of JSON, then the full report. Results are cached by URL for an hour, so a repeat comes back instantly.
# analyze a link, one JSON object per line curl -N -X POST https://halo.johanthegoat.xyz/api/analyze \ -H "Content-Type: application/json" \ -d '{"url": "https://github.com/owner/repo"}' {"type": "stage", "stage": "collecting"} {"type": "stage", "stage": "detecting"} {"type": "report", "report": { "ai_probability": 0.06, "interval": { "low": 0.01, "high": 0.35 }, "tier": "human", "top_evidence": [ ... ], "detectors": [ ... ], "limitations": [ ... ] }}
Python 3.12 and git are all it needs. SQLite and an in-process queue by default, Postgres and Redis when you scale out. Containers run as non-root on a read-only filesystem.
$ make install $ halo analyze https://github.com/owner/repo --evidence $ halo keys create --name local $ halo serve