Halo
AI provenance · pipeline 2026.10.1

A reverse GitHub for AI-written work.

Give Halo any link. It returns a calibrated estimate of how likely the work behind it was AI-generated, AI-assisted or human-written, with every claim pointed at a file, commit or line.

Public links only. Most repositories finish in under a minute.
Reads
  • repos
  • users
  • gists
  • files
  • pull requests
  • commits
  • GitLab
  • Hugging Face
  • npm
  • PyPI
  • any web page

What a tool declares is kept apart from what Halo infers.

A trailer proves a tool took part in a commit, not how much it wrote. So the estimate is built in three layers, and each one is reported on its own.

01 · Declared

Provenance

Co-Authored-By trailers, agent bot accounts, CLAUDE.md, .cursorrules, AGENTS.md and builder markers from Lovable, v0, Bolt and Replit. Anchored on emails and tool spellings, so a human co-author named Claude does not match.

02 · Inferred

18 statistical detectors

Commit timing, history mess, message style, account behavior, stylometry on tree-sitter ASTs for 10 languages, prose statistics and drift from the author's own pre-2021 writing. A detector without enough data never guesses.

03 · Fused

Calibrated estimate

Weighted log-odds learned by regularized logistic regression, calibrated with Platt or isotonic regression, with a 90 percent bootstrap interval and thresholds set for a low false-positive rate.

  1. resolving
    The URL is matched to an adapter and normalized. Private, loopback and reserved addresses are refused.
  2. collecting
    A bare, sandboxed clone reads up to 1000 commits and 4000 files. Nothing from the repository is executed.
  3. detecting
    Detectors run concurrently with a timeout each. One failure marks only that detector.
  4. reporting
    Evidence is ranked by log-odds share, and limitations are written from what was missing.

Every signal, by family.

Each detector returns a score, a confidence, an applicability flag and evidence with a pointer, plus the raw features so a reader can audit the number.

Behavioral
commit_intervals · circadian_rhythm · productivity_plausibility · history_messiness · commit_message_style
Code
ast_structure · naming_conventions · formatting_uniformity · comment_style · defensive_code · duplication
Text
llm_lexicon · text_statistics · self_baseline_drift
Account
account_behavior
Model, optional
model_code_perplexity · model_text_perplexity · model_binoculars
Learned fusion weightsShipped model v1-github. Non-negative, L2-regularized toward a prior of 0.35.
commit_message_style1.769
naming_conventions1.514
account_behavior1.064
history_messiness0.960
formatting_uniformity0.787
self_baseline_drift0.439
circadian_rhythm0.401
productivity_plausibility0.254

Seven detectors were driven to zero by the data. Reports name them in their limitations, and a target whose only applicable detectors carry no weight is reported as insufficient_data instead of a number.

Scored with the trailers stripped out.

84 real GitHub repositories from 78 owners, cross-validated in 5 folds grouped by owner. Provenance was masked and footprint files removed, so these numbers measure inference alone.

0.976AUROC
0.981PR-AUC
0.085Expected calibration error
0.016Largest AUROC drop from removing any one detector
Operating points, out-of-fold
ThresholdTierTrue positivesFalse positives
0.60mixed_or_assisted82.5%0 of 40
0.70none77.5%0 of 40
0.85mostly_ai60.0%0 of 40

Four tiers, never a verdict

  • humanBelow the 0.60 floor.
  • mixed_or_assistedAbove 0.60, or any declared tool.
  • mostly_aiPoint estimate and interval floor both clear their thresholds.
  • insufficient_dataNo weighted detector applied.

One POST, one report.

The hosted endpoint streams each stage as a line of JSON, then the full report. Results are cached by URL for an hour, so a repeat comes back instantly.

  • No sign-up for the hosted endpoint, rate-limited per address
  • Bare clones with hooks, filters and redirects disabled
  • SSRF-safe fetching, no private or reserved hosts
  • Self-host for API keys, webhooks and a durable queue
curlstreamed response
# analyze a link, one JSON object per line
curl -N -X POST https://halo.johanthegoat.xyz/api/analyze \
  -H "Content-Type: application/json" \
  -d '{"url": "https://github.com/owner/repo"}'

{"type": "stage", "stage": "collecting"}
{"type": "stage", "stage": "detecting"}
{"type": "report", "report": {
  "ai_probability": 0.06,
  "interval": { "low": 0.01, "high": 0.35 },
  "tier": "human",
  "top_evidence": [ ... ],
  "detectors": [ ... ],
  "limitations": [ ... ]
}}

Run it on your own machine.

Python 3.12 and git are all it needs. SQLite and an in-process queue by default, Postgres and Redis when you scale out. Containers run as non-root on a read-only filesystem.

$ make install
$ halo analyze https://github.com/owner/repo --evidence
$ halo keys create --name local
$ halo serve