Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

ballotbench

ballotbench is a hackathon portal you run yourself. Teams submit projects, judges score them against a rubric the organizer weights, and the organizer gets a ranking that takes each judge’s habits out and says honestly how sure it is. It starts with one docker compose up and needs no network after that.

I built it for Dogfood 2026, a hackathon where the thing you build is a hackathon platform, and the winner’s portal gets used to run the next events. That shaped every decision: it has to come up on a volunteer’s laptop, and its results have to survive a team asking “why did we come fourth?”

What makes a ranking defensible

Three things, which the rest of this book keeps coming back to:

  1. Nobody can see what they shouldn’t, or change what they shouldn’t. Judges can’t read each other’s scores; nobody can edit a project after the deadline. These rules live in the database, not only in the pages, so they hold for every way in: the web pages, the API, the admin, a shell.
  2. The maths is written down and tested. Judges score differently; the calibration model that corrects for it is described in full, and its guarantees (shifting or stretching one judge’s scores changes nothing, a judge who gives everyone the same score counts as absent) are checked by tests and by a command you can run.
  3. Everything leaves a trail. Every change writes a row to an audit log that the database won’t let anyone edit, chained by hashes so a quiet edit by a database superuser still shows.

Where to start

The source is at https://github.com/keirsalterego/ballotbench, MIT licensed.

This book

This book is published at https://keirsalterego.github.io/ballotbench/, rebuilt whenever the docs change on main. Its source is in docs/, and the chapters on judging, architecture and the data model include the repository’s top-level documents directly, so there is one copy of each. To read it locally with mdBook 0.5:

mdbook serve docs --open

Run it in five minutes

You need Docker with the Compose plugin. Nothing else: no Python, no Postgres, no accounts anywhere.

git clone https://github.com/keirsalterego/ballotbench.git
cd ballotbench
docker compose up

The first run builds the image, starts Postgres, creates the tables, loads the shared Dogfood fixture event and prints something like:

fixture event sample-hack-2026: 8 tracks, 30 judges, 40 teams, 91 members, 41 projects, 126 reviews, 1 duplicate
seeded. test logins:
  organizer    Authorization: Bearer bb_demo_organizer_5c1e0a
  judge_a      Authorization: Bearer bb_demo_judge_a_8d24f1
  judge_b      Authorization: Bearer bb_demo_judge_b_3a9e77
  participant  Authorization: Bearer bb_demo_participant_61b0c4

Open http://localhost:8080. You’re looking at the public gallery.

Sign in

Every demo account has the password ballotbench-demo:

WhoEmail
Organizerorganizer@ballotbench.local
Judge Adiego.herrera@example.org
Judge Bines.rocha@example.org
Participantpriya1@example.org
Site adminadmin@ballotbench.local

The bearer tokens are for scripts and the API: pass them as an Authorization header.

Two events are waiting

  • Sample Hack 2026 is the fixture: 41 projects, 126 reviews, on its real dates, so submissions closed in March 2026. Sign in as the organizer, open it, and go straight to Calibration and results.
  • Demo Hack (open) is empty and open for submissions for two weeks from your first boot. Use it to walk through an event yourself; the tour does exactly that.

Check what it claims

python3 run.py .dogfood.toml                                   # the official Dogfood checker (T1, T2)
python3 scripts/check_t3_t4.py .dogfood.toml                   # the same style of checks for T3 and T4
sh scripts/isolation_curl.sh                                   # tries to reach what it shouldn't
docker compose exec web python manage.py normalization_proof   # the judging maths, recomputed
docker compose exec web python manage.py verify_audit          # the audit log's hash chain

With the network off

The simplest proof: once the images are built, switch off your Wi-Fi (or pull the cable) and run docker compose up. The portal comes up, seeded, and http://localhost:8080 works, because a port on your own machine needs no outside network.

For a stricter proof, with no route out of the containers at all, there’s an override. A network with no way out can’t publish a port to your machine either, so you check it from inside the container, which is what CI does on every push:

docker compose down -v
docker compose -f docker-compose.yml -f docker-compose.offline.yml up -d --wait
docker compose exec web python -c "import urllib.request as u; print(u.urlopen('http://localhost:8080/projects').status)"
docker compose down && docker compose up -d --force-recreate      # back to normal

Start over

docker compose down -v deletes the database volume. The next up seeds a fresh copy.

A whole event, start to finish

This walks the open demo event through its life: set up, teams, submissions, the deadline, judging, calibration, results. It’s also the script of the five-minute demo video. Keep two browser windows open (a normal one and a private one) so you can be two people at once.

1. The organizer sets it up

Sign in as organizer@ballotbench.local, open My events, then Manage next to Demo Hack (open).

  • Dates and settings. Four windows: submissions, judging, voting and (implicitly) results. k, reviews per project, defaults to 3.
  • Tracks and prizes. Add a track called Hardware; add a prize for it.
  • Rubric. Three criteria to start: functionality, quality, innovation, each 1 to 5, weight 1. Give functionality weight 2. The weights can change freely until the first score comes in; after that the database refuses, so nobody can tune weights after seeing who they favour.
  • Invite a judge. Under Invite judges and organizers, pick Judge, tick a track, and make a link. It’s shown once and works once. The portal sends no email; send the link however you like.

2. Participants form a team

In the private window, create an account, then My events → join as a participant on the demo event.

  • Create team, then Make an invite link. Open it signed in as someone else and they join. The link then stops working; a team can’t grow past the event’s size limit, even if two people accept at the same moment.
  • Start your project. Save it as a draft: only the team and the organizers can see it. Submit it and it appears in the public gallery. You can keep editing until the deadline.

3. The deadline

As the organizer, move Submissions close to a minute from now and save. Wait for it. Back in the participant’s window, try to save the project: the answer is That window is closed, with the time. The same happens through the API (a 409 with the close time) and in the admin. Behind all three, a trigger in the database refuses the write on its own clock, so even a bug in the app couldn’t let a late edit through.

4. Judging

As the organizer, open Hand out reviews. The preview shows which judge gets which project: every project gets k reviews, the idlest judge picks next, nobody reviews their own team or outside their tracks. Apply it.

Sign in as the judge. Your reviews lists only your projects and how long they’ll take. Each has a scoresheet: one bubble per score per criterion, a comment only the organizers see, Save draft and Submit review (after submitting, only Resubmit). If you know the team, I have a conflict steps you aside; the organizer sees the project needs another judge.

Meanwhile, the organizer’s Progress page shows each judge’s assigned, started and finished counts and which projects are short of reviews. Keep this page up to date refreshes it every 20 seconds.

5. Calibration

Open Calibration and results and run it. For the fixture event this is where it gets interesting:

  • jdg_07 gave every project 4 on everything. They’re listed as gave every project the same score and carry no weight.
  • jdg_01 and jdg_23 wrote one review each: nothing to compare with, so no weight either.
  • The ranking shows raw rank, calibrated rank, and the range each rank could plausibly be. On the fixture these ranges are wide, and the page says why: the judges agree no more than chance would.

Reading the calibration page goes through it line by line.

6. Publish

Publish results freezes the ranking to this calibration run. Until then, the results page and API answer 404 to everyone but the organizers. The public page shows each project’s calibrated score, its plausible rank range, and the fingerprint of the scores it came from.

7. A community vote

In Settings, set Voting mode to signed-in accounts and put the voting window after submissions close (the portal refuses it earlier). While it’s open, Community vote shows how many ballots are in and who cast them, with flags for crowded networks, brand-new accounts and identical ballots, but not the per-project tallies: those appear once voting closes.

A voter opens Vote on the event: the projects come in an order made for them, they spread 25 credits (n votes on a project cost n²), and they can’t vote for their own team. An account has to confirm its email address first: the link is sent at sign-up, and with the network off it lands in the outbox, which the site admin reads under Admin → Outbound emails.

Results can’t be published while the vote is open. Close it, publish, and the public results show the community votes next to the judges’ ranking.

8. Export

Exports has a CSV for every stage: registrations, teams, projects, assignments, every score, the calibrated results, each judge’s calibration, and the audit log with its hash chain. Every download is itself in the audit log.

The five-minute demo

The automatic version

scripts/demo.sh runs a whole event against the portal and narrates it, printing the page to show in the browser at each step:

From a fresh clone, nothing else to set up:

git clone https://github.com/keirsalterego/ballotbench.git && cd ballotbench
sh scripts/demo.sh                     # builds and starts the portal if it isn't running, then runs
sh scripts/demo.sh --fresh --pause     # empty database first, then wait for Enter before each step
BALLOTBENCH_PORT=9000 sh scripts/demo.sh   # if port 8080 is taken

It creates a new event each run (Demo Day plus the time), so it can run again without a reset. In twelve steps:

  1. An organizer creates the event, with a track, a prize and a weighted rubric.
  2. Four judges accept single-use invites.
  3. Four teams sign up and submit; one invites a teammate.
  4. The deadline passes: a late edit is refused by the API and the page.
  5. The organizer hands out reviews.
  6. Four judges score: harsh, generous, steady, and one who gives everything a 3.
  7. Isolation attacks between judges and teams are refused.
  8. Calibration flags the constant judge and ranks with honest ranges.
  9. A community vote with a quadratic budget, each voter in their own order, and publishing refused while it’s open.
  10. Publishing, the signed results verified, a doctored copy refused.
  11. A team reads why it placed where it did; another team can’t.
  12. The exports and the audit trail.

It needs only Python 3 (standard library) and, for --fresh, Docker.

By hand

A shot list for recording one full event lifecycle (create, submit, judge, publish) in five minutes. It uses both seeded events: the open Demo Hack for the live lifecycle, and Sample Hack 2026 (the fixture, 126 real reviews) for judging maths worth showing.

Before recording:

docker compose down -v && docker compose up -d --build --wait

Open three browser windows (a normal one and two private ones) so you can be three people at once. Every demo password is ballotbench-demo.

TimeWhoWhat to doWhat to say
0:00anyoneOpen http://localhost:8080.One docker compose up, no network needed, seeded with the shared Dogfood fixture.
0:15organizer@ballotbench.localMy events → New event, or open Demo Hack (open) → Manage. Show dates, add a track, set functionality weight 2.Organizers set windows, tracks, prizes and a weighted rubric. The rubric freezes at the first score.
0:45new account (private window)Create an account, join as a participant on Demo Hack, Create team, Make an invite link.Single-use invite links, stored only as hashes.
1:05sameStart your project, fill title and repo, Submit. Show it in the gallery, then open Edit again and leave that tab open.Drafts are private; submitted projects are public.
1:25organizerSettings: move Submissions close to a minute ago, save.Watch the deadline hold on every path.
1:35participantIn the edit tab you left open, press Save changes: That window is closed. Reload the event page: the project is now read-only. Then in a terminal, the fixture event, which closed in March: curl -X POST -H "Authorization: Bearer bb_demo_participant_61b0c4" -H "Content-Type: application/json" -d '{"title":"late"}' localhost:8080/api/events/sample-hack-2026/projects → 409.The page, the API and a database trigger on the database’s own clock all refuse; even a bug in the app couldn’t let a late edit through.
1:55terminalpython3 run.py .dogfood.tomlThe official checker: T1 and T2 verified.
2:05terminalsh scripts/isolation_curl.sh (let it scroll)94 attempts, each answered with exactly the status it must get: 78 refused, 16 allowed (public pages and your own data).
2:15diego.herrera@example.org (judge)Your reviews (in the header) → open a project: the scoresheet with ballot bubbles.Judges see only their own assignments; another judge’s scoresheet is a 404, naming another judge in the API is a 403.
2:35organizerSample Hack 2026 → Manage → Hand out reviews: the preview proposes 8 top-ups.The fixture’s two unfinished batches: the planner finds them and balances load.
2:50organizerCalibration and results → Run calibration. Scroll: flagged judges, does one judge decide the podium?, the ranking with could be ranges and the pairwise column.The fixture’s constant judge (jdg_07, listed by name as Iva Petrova) gave every project 4: weight zero, exactly as if absent. Every rank comes with an honest range; on this data the judges agree no more than chance, and the page says so.
3:35organizerPublish results. Open the public results page; click signed copy; paste it into /verify: Valid.Published results are frozen to one run, fingerprinted, and signed with Ed25519.
3:55priya1@example.org (team NorthKiln)Results → How your project was scored (Glass Signal, third).“Why did we come third?” answered: each review, the judge’s habit, how much it counted. Judges stay anonymous.
4:15organizerAudit log: filter by results. Then Exports: download scores.csv.Every change is in a hash-chained log the database won’t let anyone edit. Every stage exports as CSV.
4:35terminaldocker compose exec web python manage.py verify_audit and ... normalization_proof | tail -8The chain checks out; the maths claims are recomputed live and hold.
4:50Open the book: https://keirsalterego.github.io/ballotbench/ (it redirects to the custom domain).MIT licensed; everything above is documented for a stranger.

If there’s time to spare, show the community vote on Demo Hack: in Settings set Voting mode to signed-in accounts with a window that’s open now, then vote as priya (confirmed demo account) and show the quadratic budget and the organizer’s abuse panel.

For organizers

You organize an event if you created it, or someone sent you an organizer invitation for it. Site admins can do everything organizers can, in every event.

Creating an event

My events → New event. Anyone who already organizes an event, and any site admin, can create one; you become its first organizer. Every event starts with a three-criterion rubric you can change.

The slug is part of every URL and can’t change later. All times are UTC.

The windows

WindowWhat it controls
Submissions open → closeTeams can form, and projects can be created, edited and submitted. Enforced by the database clock.
Judging open → closeJudges can save and submit reviews.
Voting open → closePublic voting, if you turn it on. Results can’t be published while it’s open.

Moving a window takes effect immediately: extending the deadline reopens editing at once.

Tracks, prizes and the rubric

Tracks group projects and limit which judges see what: a judge with tracks only reviews projects in those tracks. A judge with no tracks reviews anything.

The rubric is a list of criteria, each with a weight and a range. A review’s score is the weighted mean of its criteria, each scaled to 0..1 by its own range. Change it freely until the first score arrives; after that the database freezes it.

Judges and co-organizers

Invite them with a link from the settings page. Links are single use and last seven days. A judge invitation can carry tracks. Nobody can be both a judge and on a team in the same event: the database refuses either order.

Handing out reviews

Hand out reviews shows a plan before it changes anything. It tops every submitted project up to k reviews, gives each new review to the idlest judge who may take it, and adds a few bridging reviews if some judges share no projects with the rest (calibration needs them connected). Projects no eligible judge is left for are listed, so you know to invite someone.

You can also give one project to one judge by hand on the Progress page, and take back a review nobody has started.

Watching progress

Progress lists each judge’s assigned, started, finished and stepped-aside counts, and each project’s finished reviews, fewest first. Projects below k are flagged low coverage.

The community vote

Turn it on in Settings with Voting mode: signed-in accounts (anyone with an account votes once) or confirmed email addresses (anyone votes once per inbox, after opening a link we mail them). Vote credits is each voter’s budget: n votes for one project cost n² credits. Send people to /events/<slug>/vote; the gallery and every project page link there too.

With the network off, confirmation links can’t leave the box. They land in Outbound emails in the admin, where a site admin can read them.

Community vote in the organizer menu shows how many ballots are in while voting is open, and the tallies once it closes (only organizers see them until you publish results, which you can’t do while voting is open), and a list of things worth a second look: several voters on one network, accounts made just before their ballot, identical ballots, and sign-ups refused because another spelling of the same inbox had already voted. None of these void anything by themselves. If you decide a ballot is fake, void it with a reason; it leaves the tallies for good and the audit log records it. There’s no undo: voiding a ballot and counting it again would show you what that voter chose.

Duplicates

When a project has the same repository and title as an earlier one (or the same team resubmits the same repository or title), it’s flagged and left out of the rankings. Only submitted projects count, and the one submitted later is the copy. Duplicates lets you confirm it or put it back; a project you put back is never flagged again.

Calibration and results

Run calibration as often as you like; each run is kept. Publishing freezes the public results to one run. You can hide them again. See Reading the calibration page.

The audit log

Audit log lists every change in the event: who, when, from which address, and the before and after values. Filter by kind of change or by person. Rows can’t be edited or deleted, even by the admin; verify_audit checks the hash chain.

Exports

Every stage as CSV. Cells that a spreadsheet would run as a formula (a title starting with =, say) are prefixed with a quote, so opening an export is safe.

For participants

Joining

Create an account (Create an account, top right), then on My events choose join as a participant next to the event. If a teammate sent you an invite link, just open it and sign in; it adds you to their team.

Your team

You can be on one team per event. Start one from the event page, then Make an invite link for each teammate. Each link works once, and stops working after 72 hours or when submissions close, whichever comes first. You can revoke a link you haven’t used. The event sets the largest team size.

Your project

A team usually submits one project, but may submit more (the fixture has a team with two); each is judged on its own.

Start your project and fill in what you have: title, a one-line tagline, a summary, the description, links to the code and a demo, tags and a track.

  • Save draft keeps it private to your team and the organizers.
  • Submit puts it in the public gallery.
  • After that, Save changes updates it; it stays submitted and public. You can keep editing until the deadline.
  • A draft you don’t want can be deleted from its edit page (Delete this draft). A submitted project can’t be deleted by the team; ask the organizers.

Starting a second project with the same title as one your team already has takes you to the existing one instead, so a double click doesn’t leave two copies.

The deadline

At the close time, editing stops: the edit page shows your project read-only, and the page, the API and the database all refuse changes, on the server’s clock, not yours. If the organizers extend the deadline, editing reopens straight away.

Duplicates

If your project looks like one submitted earlier (same repository and title), it’s flagged for the organizers when you submit it, and left out of the rankings until they decide. Only the later submission is ever flagged, and once an organizer says it’s a different project, it stays that way. This catches accidental double submissions; if yours is flagged by mistake, tell the organizers.

For judges

An organizer sends you an invitation link. Open it, sign in or create an account, and accept. Your reviews then lists the projects assigned to you, and how long the rest will take at about ten minutes each.

Scoring

Each project has a scoresheet: the project’s description and links, then one row of bubbles per criterion. Pick a number in each row. The weight of each criterion is shown next to its name.

  • Save draft keeps your scores without submitting them.
  • Submit review records them. After that the sheet offers only Resubmit review: you can change your scores and resubmit until judging closes or the results are published, but a submitted review never goes back to being a draft.
  • The comment goes to the organizers. Other judges never see it.

What you can and can’t see

You see your own assignments and your own scores, nothing else. Another judge’s review isn’t hidden on a page: it’s not reachable at all. The server answers “not found” for it, whatever URL you type or tool you use.

Conflicts

If you know the team, or have any other reason you shouldn’t judge a project, open it before you submit a review and choose I have a conflict, giving the reason. You’re taken off it, the organizers see it needs another judge, and it’s recorded in the audit log with your reason. You can’t be assigned a project from a team you’re on in the first place.

How your scores are used

Everyone reads a rubric a little differently: some judges are generous, some harsh, some use the whole scale and some stay near the middle. The calibration learns your habits from the projects you share with other judges and takes them out, so it doesn’t matter whether your 4 is someone else’s 3. If you give every project the same score, your scores carry no weight: they say nothing about which project is better. The method explains it all.

Assignment, scoring and calibration

This chapter is the repository’s JUDGING.md, included here as it stands so the two can’t drift apart. It covers how reviews are handed out, how a scoresheet becomes one number, the model that takes each judge’s habits out of the ranking, what that model guarantees and where it stops. Every figure in it comes from manage.py normalization_proof, which you can run yourself. If you want to know what the calibration page shows you rather than how it’s computed, read Reading the calibration page next.

How ballotbench hands out reviews, turns rubric scores into a ranking, and why I think the ranking can be defended to a team that didn’t win. Every number on this page comes from manage.py normalization_proof, which recomputes it from the database; the output is in section 8.

1. Assigning reviews

portal/assignment.py is a pure function over plain records, so it’s tested on thousands of random events without a database (tests/test_assignment.py).

  1. The idlest judge picks next. Of the judges who can still help, the one with the fewest reviews takes a project. That keeps loads within one review of each other wherever eligibility allows.
  2. They take the neediest project they may review. Fewest reviews so far, then fewest judges left who could take it, so scarce judges go where only they can help.
  3. Eligibility is enforced in the planner and again by a database trigger: never a project from your own team, only your tracks (a judge with no tracks takes any), never a project you already hold or stepped aside from.
  4. Repair. Greedy can corner itself: at the end, the idle judges may already hold the last projects that need someone. A repair pass moves one of the plan’s new reviews from the busiest judge to the idlest eligible one until no move narrows the gap. Existing reviews never move.
  5. Connectivity. Calibration compares judges through the projects they share. If the judge-project graph splits into groups that share nothing, their scales can’t be compared, so the planner adds bridging reviews (union-find over the graph) and the page says how many it added.

The organizer previews the plan, then applies it; on apply it’s recomputed from the database, never taken from the form. Manual changes are allowed, but a review that has been started can’t be taken back, since that would quietly delete a judge’s scores. A judge who spots a conflict steps aside; the project shows up as needing a top-up.

On the fixture, the plan for k = 3 is exactly the eight top-ups the fixture’s two unfinished batches call for: prj_10, prj_15, prj_18, prj_19, prj_24, prj_29, prj_39, prj_40 each have two reviews. The progress dashboard flags them as low coverage until the new reviews come in.

2. A review’s score

Each criterion has a weight and a range (1 to 5 by default). A review’s score is the weighted mean of its criteria, each first mapped to 0..1 by its own range, so a 1-10 criterion and a 1-5 criterion count by their weights, not by the width of their scales:

score = Σ_c w_c · (x_c − min_c) / (max_c − min_c)  /  Σ_c w_c

A review missing a criterion has no score and isn’t used. The fixture’s scale isn’t stated; its values run 2 to 5, so I assume 1 to 5.

The rubric freezes at the first score. A trigger refuses any change to criteria, weights or ranges once the event has a score. Weights tuned after reading the scores are a way to pick a winner.

3. Why not just average, or z-score

  • Raw means reward drawing a lenient judge. With three reviews per project, one generous judge moves a project a long way.
  • Per-judge z-scores fix leniency but assume every judge saw an average batch. A judge who only saw the five best projects gets the best of those pushed down to the middle. They also divide by zero for a judge who gives the same score every time, which the fixture has (jdg_07).

4. The model

Each review’s score y is modelled as

y_ij = a_j + s_j · q_i + e_ij,     e_ij ~ N(0, v_j),     q_i ~ N(0, 1)
  • q_i: project i’s quality, the thing we want.
  • a_j: judge j’s leniency (offset).
  • s_j: how strongly their scores follow quality (scale). A judge who separates good from weak sharply has a large s_j.
  • v_j: how noisy they are.

The fit (portal/calibration.py) alternates two least-squares steps until q stops moving (tolerance 1e-12):

q_i  = Σ_j s_j (y_ij − a_j) / v_j   /   (1 + Σ_j s_j² / v_j)        then centre and scale q to mean 0, sd 1
s_j  = S_xy / (S_xx + 1),   a_j = ȳ_j − s_j · q̄_j                  ridge least squares on the judge's own projects, s_j ≥ 0
v_j  = (3 · var(y_j)/2 + RSS_j) / (3 + n_j)                         empirical-Bayes shrink of the noise
  • The q_i step is the posterior mean under the N(0, 1) prior: a project seen by few or noisy judges is pulled towards the middle instead of trusted blindly. That’s the answer to the unfinished batches.
  • A judge’s leniency is anchored by the projects they share with other judges, not by their own batch, which is what z-scores get wrong.
  • The noise prior stops a judge with three reviews from being fitted exactly and then trusted infinitely.
  • The ridge on s_j (the + 1) is there because my first version had none, and on the fixture it didn’t converge: a judge whose few projects landed close together in q got a huge scale, which dragged those projects closer, and so on. The penalty is in q’s units, which have no scale of their own, so it doesn’t break the invariances below.

This is the reviewer-calibration model used for NeurIPS reviewing (Lawrence, 2014; Ge, Welling & Ghahramani), fitted by alternating least squares rather than full Bayesian inference, which is plenty for tens of judges.

Judges it can’t learn from

These get s_j = 0. Because every term a judge contributes to q_i is multiplied by s_j, a judge with s_j = 0 drops out exactly, as if their reviews weren’t there. Raw means still include them; the calibration page lists them with the reason.

FlagMeaningOn the fixture
constantevery review the same scorejdg_07 (4 on every criterion, 3 reviews)
single_reviewone review: nothing to compare it withjdg_01, jdg_23
discordanttheir scores fall, or don’t rise, as everyone else’s rise on the projects they share9 judges

What it guarantees (tested)

  • Shift one judge (add a constant to all their scores): every q is unchanged, to 1e-15.
  • Stretch one judge (multiply their scores by a positive factor): same.
  • Shift and stretch every judge at once, each differently: same.
  • Remove the constant judge: every q is unchanged, exactly 0.

These hold because every step is equivariant: a judge’s a_j, s_j and v_j absorb any affine change of their own scores, and the ridge and priors are either in q’s units or built from the judge’s own spread.

5. How sure is the ranking?

The model’s own standard error, 1/√(1 + Σ s_j²/v_j), assumes each judge’s offset and scale are known exactly. With three or four reviews per judge they aren’t. So each run also bootstraps: resample each project’s reviews with replacement, refit, 200 times, and record where each project lands. The pages show the 90% range as “could be 4 to 12”. Where two projects’ ranges overlap, the judges couldn’t really tell them apart, and the page says so rather than printing a confident third decimal.

6. Does the data have a signal at all?

Each run also asks whether the judges agree about which projects are better more than chance would: it compares the spread of the project means with the spread after shuffling every score across the same judge-project slots (2000 shuffles). If they don’t, no method can produce a meaningful ranking, and the calibration page says so in plain words.

On the fixture they don’t (p ≈ 0.70). The fixture’s scores look like independent draws: projects’ means spread no more than shuffled scores do, and nine judges’ scores run against the consensus. So the honest result on the fixture is: the constant and single-review judges are handled, the duplicate is out, the ranking is computed, and almost every rank’s interval is wide (median 27 places). I’d rather show that than a crisp ranking that is noise.

7. The fixture’s traps

TrapWhat ballotbench doesWhere you see it
jdg_07 scores 4 on everythingflagged constant, weight exactly 0calibration page, judges.csv, the proof
jdg_01 (and jdg_23) have one reviewflagged single_review, weight 0same
two unfinished batches: 8 projects with 2 reviewsshrunk towards the middle, wider intervals, “low coverage” flag, 8 top-ups proposed. One of them, prj_24, was reviewed only by two judges the model ignores (both discordant), so it has nothing to be ranked by and is listed as left out until its top-up reviews arriveprogress dashboard, assignment preview, calibration page
prj_41 repeats prj_07 (same team, title, repo, submitted later)detected at import, duplicate_of set, out of the rankings until an organizer decides; its reviews still calibrate its judgesduplicates page, gallery badge, audit log
judge load 1 to 11noise shrinkage trusts busy judges more; new work goes to the idlestjudge table

8. The proof on the fixture

docker compose exec web python manage.py normalization_proof, on a fresh seed:

Normalization proof: Sample Hack 2026, 126 reviews, 30 judges, 41 projects
Model: y = a_j + s_j*q_i + e, fitted in 210 iterations (converged), judge-project graph in 1 component(s).
Judges carrying no weight (s_j = 0):
  jdg_07    3 review(s)  constant
  jdg_04    4 review(s)  discordant
  jdg_08    3 review(s)  discordant
  jdg_10    3 review(s)  discordant
  jdg_14    3 review(s)  discordant
  jdg_17    2 review(s)  discordant
  jdg_18    3 review(s)  discordant
  jdg_19    4 review(s)  discordant
  jdg_21    4 review(s)  discordant
  jdg_27    2 review(s)  discordant
  jdg_01    1 review(s)  single_review
  jdg_23    1 review(s)  single_review
  the other 18 judges: ok
Agreement: variance of project means 0.0080, 0.0089 on average when every score is shuffled across the same slots (permutation p = 0.698, 2000 shuffles).
  The judges agree on which projects are better no more than chance would. Calibration takes out
  judge habits; it can't create a signal the scores don't hold, so most rank moves below are noise.
Left out of the ranking: prj_41 (Dry Harbour) repeats prj_07; its 4 reviews still count towards calibrating their judges.
Left out of the ranking: prj_24 (Glass Beacon): all 2 of its reviewers carry no weight, so there's nothing to rank it by.
rank  could be  raw  move  project  title             n raw mean calibrated
   1      1-39   31   +30  prj_07   Dry Harbour       5    0.583      0.832
   2      1-33    5    +3  prj_37   Salt Loom         4    0.771      0.804
   3      2-36   22   +19  prj_01   Glass Signal      3    0.611      0.774
   4      2-37    9    +5  prj_08   North Drift       5    0.700      0.762
   5      1-16    1    -4  prj_11   Salt Ledger       4    0.833      0.749
   6      2-12    2    -4  prj_34   Iron Switch       3    0.833      0.746
   7      3-36   15    +8  prj_09   Hollow Signal     3    0.639      0.741
   8      4-36   25   +17  prj_12   Open Beacon       3    0.611      0.739
   9      4-39   26   +17  prj_27   Flat Thread       3    0.611      0.732
  10      2-29    4    -6  prj_25   Dry Relay         3    0.778      0.716
  11      3-16    7    -4  prj_33   Slow Trail        3    0.750      0.709
  12      4-31    8    -4  prj_21   Copper Kiln       3    0.722      0.708
  13      6-34   18    +5  prj_02   Small Meadow      3    0.639      0.686
  14      8-28   17    +3  prj_31   Salt Ferry        3    0.639      0.684
  15      3-35   10    -5  prj_04   Green Switch      3    0.694      0.675
  16      5-20    3   -13  prj_10   Still Beacon      2    0.792      0.666
  17      3-38   21    +4  prj_35   Warm Beacon       5    0.617      0.652
  18      3-36   30   +12  prj_29   Flat Relay        2    0.583      0.651
  19      5-32   12    -7  prj_15   Copper Orbit      2    0.667      0.651
  20     11-37   28    +8  prj_03   Deep Compass      3    0.583      0.634
  21      1-29    6   -15  prj_16   Salt Kiln         3    0.750      0.624
  22     11-29   14    -8  prj_36   Salt Drift        3    0.667      0.615
  23     18-29   13   -10  prj_19   Small Relay       2    0.667      0.589
  24      4-27   16    -8  prj_17   Small Loom        3    0.639      0.589
  25     12-33   27    +2  prj_14   Green Lantern     5    0.600      0.582
  26      5-36   19    -7  prj_18   Open Kiln         2    0.625      0.580
  27     21-37   23    -4  prj_28   Flat Meadow       3    0.611      0.579
  28      1-33   20    -8  prj_39   Paper Anchor      2    0.625      0.571
  29      2-35   11   -18  prj_38   Deep Beacon       3    0.694      0.567
  30     21-37   37    +7  prj_40   Slow Loom         2    0.500      0.564
  31      4-38   36    +5  prj_30   Paper Harbour     3    0.528      0.559
  32     18-37   34    +2  prj_06   Dry Compass       3    0.528      0.558
  33     17-39   35    +2  prj_22   Dry Bridge        3    0.528      0.556
  34      4-38   33    -1  prj_26   Amber Hours       3    0.556      0.544
  35      5-38   24   -11  prj_32   Loud Ledger       3    0.611      0.529
  36     12-39   29    -7  prj_13   Quiet Anchor      3    0.583      0.515
  37     14-39   32    -5  prj_20   Paper Thread      3    0.556      0.504
  38     28-39   38        prj_05   North Compass     3    0.472      0.500
  39     24-39   39        prj_23   Slow Quarry       3    0.472      0.490

37 of 39 projects change rank. Kendall's tau, calibrated vs raw: 0.457
'could be' is a 90% bootstrap interval (each project's reviews resampled, 200 refits). Median width 27 places: on these scores most ranks are not distinguishable.
Invariance checks on these scores (max |change in q| over all projects):
  PASS  jdg_24 adds 0.3 to every score: 1.1e-15, ranking identical
  PASS  jdg_24 doubles every score: 0.0e+00, ranking identical
  PASS  every judge gets a random shift and stretch: 3.1e-15, ranking identical
  PASS  constant judge(s) jdg_07 removed: 0.0e+00, ranking identical

Kingmaker check: leaving out one judge at a time, 9 of 30 judges' absence would change who is in the top 3:
  without jdg_11: prj_35 in, prj_01 out
  without jdg_02: prj_16 in, prj_01 out
  without jdg_20: prj_08 in, prj_01 out
  without jdg_22: prj_08 in, prj_37 out
  without jdg_26: prj_39 in, prj_07 out
  without jdg_28: prj_08 in, prj_01 out
  without jdg_03: prj_08 in, prj_01 out
  without jdg_05: prj_08 in, prj_01 out
  without jdg_06: prj_08 in, prj_01 out

Second opinion: the rubric read as pairwise picks (Bradley-Terry, portal/pairwise.py):
  254 picks; no picks from jdg_01, jdg_07, jdg_23 (one review, or every pair tied)
  Kendall's tau with the calibrated ranking 0.649, with raw means 0.630
  PASS  jdg_24's scores squared (not a shift or stretch): picks identical; the calibration's q moves by up to 0.04, since it only undoes linear habits

Benchmark on synthetic events with a known true order (40 projects, 12 judges, 3 reviews each,
one harsh judge who only sees the strongest projects). Kendall's tau against the truth:
  calibration 0.830   raw means 0.677   per-judge z-scores 0.742   pairwise picks 0.799   (mean of 20 events)
  calibration beats raw means in 20 of 20, z-scores in 20 of 20
Every invariance check holds.

The benchmark events are synthetic because the fixture has no ground truth. Each has 40 projects, 12 judges with their own leniency, scale and noise, and one harsh judge who only sees the eight strongest projects: the case that breaks z-scores. The model recovers the true order better than raw means and better than z-scores in every one of the 20 events, and on average better than the pairwise second opinion, which throws away how far apart a judge put two projects.

9. A second opinion: pairwise picks

The calibration undoes linear habits. A judge who squashes only the top of the scale is non-linear, and it can’t. So the calibration page shows a second ranking next to it that no way of using the scale can move.

Every judge who scored two projects differently has, in effect, picked the better one. portal/pairwise.py collects those picks across all judges and fits a Bradley-Terry model, P(A beats B) = p_A / (p_A + p_B), with Hunter’s MM iteration and one virtual win and loss per project against a reference, a weak prior that keeps an unbeaten project finite. Only the order of each judge’s own scores enters, so any increasing transformation of any judge’s scores leaves it unchanged; the proof squares one judge’s scores to show it. The constant judge and the single-review judges make no picks at all.

What it gives up is magnitude: “A slightly better than B” and “A far better than B” are the same pick. That’s why it’s a second opinion rather than the ranking. Where the two agree, trust the rank more; where they disagree, the project’s rank depends on how you read the judges’ scales, and its “could be” range will usually be wide as well.

10. Does one judge decide the podium?

Before publishing, the calibration page runs the whole fit again once per judge, each time without that judge’s reviews, and lists every judge whose absence would change who is in the top three (or the top N, for N prizes). It takes about a second on the fixture.

This finds the case the other checks can’t: a judge the model trusts, whose scores agree with everyone else’s on most projects, lifting one borderline project onto the podium. A lone contrarian is not a kingmaker: a judge whose scores run against the consensus is flagged discordant and carries no weight, so leaving them out changes nothing (tests/test_calibration.py plants both and checks each). A name on the list isn’t an accusation; it says the podium rests on one person, which is where a second look, or one more review, is worth it.

On the fixture, nine judges’ absence would each swap one podium place, consistent with scores that carry no real signal.

11. Spending the next reviews where they matter

Once a calibration has run, the assignment page lists the projects whose plausible rank range crosses the prize line (the number of prizes, or the top three): projects that could end up either side of it. Give each of these one more review adds exactly one review per contested project, from the idlest eligible judge. Spreading extra reviews evenly spends most of them on projects that can’t win and can’t miss; these are the ones where another opinion can change who wins. Run calibration again afterwards and the ranges narrow where it counted.

12. Explaining a rank to the team

After publication, each team can open How your project was scored: every review of their project with the judge anonymized (“Judge 2”, shuffled per project), what that judge gave them, what that judge gives a typical project (their offset a_j), the difference, and the share of the final score that review carried ((s_j²/v_j) / (1 + Σ s²/v)), with the pull towards the middle as its own row. A judge who counted for nothing says why in words: gave every project the same score, wrote one review, or scored against the consensus. It shows weighted totals only, no per-criterion scores, comments or names, and warns if scores changed after the published run.

13. Publishing results

Results are hidden from everyone but the event’s organizers until published, in the pages and the API (404 before then). Publishing freezes the result to one calibration run. Each run stores the SHA-256 of exactly the scores it read, in a canonical order, so anyone with the scores export can recompute it and check the published ranking came from those scores. Later runs change nothing public until someone publishes again, and every run and publication is in the audit log.

The published ranking is also available as one signed document, /api/events/<slug>/results/signed: the event, the run, the scores’ digest and every rank with its “could be” range, signed with the portal’s Ed25519 key. Anyone who saved a copy on results day can check it at /verify, or offline with the public key at /.well-known/ballotbench-signing-key, and hold the portal to it if the page ever says something else.

14. Known limits

  • Linear judges only. The model corrects a judge who is lenient or who spreads scores widely. It can’t correct one who only compresses the top of the scale, or who uses the scale differently for different criteria (Wang & Shah, “Your 2 is my 1”).
  • Collusion. Two judges who agree to push a project look like two judges who agree. Nothing statistical separates them; conflict-of-interest rules and the audit log are the defence.
  • Thin data. With three reviews per project and a few per judge, ranks are uncertain. The intervals say so; they don’t fix it. More reviews per project do.
  • One number per review. Calibration works on the weighted total. A per-criterion model would need far more reviews than a hackathon has.

References

  • D. Hunter, “MM algorithms for generalized Bradley-Terry models”, Annals of Statistics, 2004.
  • N. Lawrence, “Reviewer calibration for NIPS”, 2014. https://inverseprobability.com/2014/08/02/reviewer-calibration-for-nips
  • H. Ge, M. Welling, Z. Ghahramani, “A Bayesian model for calibrating conference review scores”. https://mlg.eng.cam.ac.uk/hong/unpublished/nips-review-model.pdf
  • M. Roos, J. Rothe, B. Scheuermann, “How to calibrate the scores of biased reviewers”, AAAI 2011.
  • J. Wang, N. Shah, “Your 2 is my 1, your 3 is my 9”, 2018. https://arxiv.org/abs/1806.05085

Reading the calibration page

The calibration page is where an organizer turns a pile of reviews into a ranking, and it’s the page I’d want open when a team asks why they came fourth. It’s at Manage → Calibration and results, or /events/<slug>/manage/calibration. Only the event’s organizers and site admins can open it; anyone else gets a 403.

This chapter goes through it block by block, using the fixture event (Sample Hack 2026) as it looks straight after docker compose up. The numbers below are from a freshly seeded stack; yours will match until someone changes a score.

Running it

The button at the top says Run calibration the first time and Run it again on the current scores after that. A run reads every submitted review that has a score for every criterion; drafts and half-filled scoresheets are ignored. It fits the model, bootstraps the rank intervals (200 refits) and runs the agreement test, which takes a few seconds on the fixture.

Every run is kept, and a run is never updated. The page always shows the latest one. That matters for two reasons: you can run it as often as you like while judging is still going, and a published result can’t change underneath you, because publishing points at one run and later runs are new rows. Each run also writes a calibration.run row to the audit log with its fingerprint, the number of reviews, the agreement p-value and the judges it flagged.

The run facts

The first block is a short list of facts about the run.

LineWhat it saysOn the fixture
Whenthe time the run was made (UTC) and who made ityour run
Reviews readsubmitted, complete reviews the run used126
FingerprintSHA-256 of exactly the scores the run read397dd35f…178f4a
Judges linkedwhether every judge can be compared with every otherYes
Agreementwhether the judges agree more than chance wouldno (p = 0.70)

Fingerprint

The fingerprint is a SHA-256 over the scores the run read, written out in a fixed order. It’s there so that a ranking can be checked against the scores it claims to come from: if one score changes, the fingerprint changes. The same fingerprint is stored with the run, shown on the public results page, returned by the results API as input_digest, and written into every row of results.csv.

On a freshly seeded fixture it’s always:

397dd35f376eb052b0dfcd8fda8309f46bad4771e9caa4e78035175807178f4a

The fingerprint doesn’t prove the scores are the ones the judges meant to give. It proves the ranking came from these scores. For the first question, the audit log records every review.save and review.submit with the before and after values, and manage.py verify_audit checks nobody has edited that log. Recompute the fingerprint below shows how to check it by hand.

Judges linked

Calibration compares judges through the projects they share. If the judges split into groups that share no project, directly or through other judges, their scales can’t be put side by side: a lenient group and a harsh group look exactly like a strong batch and a weak one.

  • Yes means every judge is connected. The fixture is.
  • No is flagged, with the advice to hand out bridging reviews. Open Hand out reviews: the planner detects the split, adds bridging reviews until the groups connect, and says how many it added. Once those reviews are in, run calibration again.

Agreement

This line answers the question to ask before looking at any rank: do the judges agree about which projects are better any more than chance would? The test shuffles every score across the same judge-project slots 1000 times and counts how often a shuffle spreads the project means at least as far apart as the real scores do. That fraction is the p-value.

  • p below 0.05: “The judges agree on which projects are better far more than chance would.” There’s a real signal. The ranking still has uncertainty, which the “could be” column shows.
  • p of 0.05 or more: “The judges agree no more than chance would”, with a flag. The ranking is computed anyway, but most of its order is noise.
  • No reviews yet if the run read nothing.

The fixture reads p = 0.70 (the page and normalization_proof run the same test: 2000 shuffles, the same seed): its scores behave like independent draws.

When p is high, I would not publish a ranking to three decimals. What I’d do, roughly in this order:

  1. Treat overlapping ranks as ties. Read the “could be” column, not the rank. If you must name winners, name a group whose ranges sit clearly above the rest, and say it is a group.
  2. Get more reviews. The signal grows with reviews per project. Raise k in the event settings, hand out more reviews, then run it again.
  3. Look at the flagged judges (next section). Nine discordant judges out of thirty, as on the fixture, says the judges weren’t reading the rubric the same way, or the rubric doesn’t separate the projects.
  4. Say so when you announce. “The judges couldn’t separate projects 4 to 20” is a fair thing to tell teams, and the public results page already shows each project’s range.

Calibration can take out a judge’s leniency and scale. It can’t create agreement that isn’t in the scores, and the page says that in as many words.

Judges the model couldn’t learn from

This table lists judges whose reviews carry no weight in the calibrated scores: their fitted scale is exactly 0, which removes them from every project’s score as if they hadn’t reviewed. Their reviews still count in the raw means and in each project’s review count. Each row shows the judge’s name (or email), how many reviews they wrote, and why.

FlagThe page saysWhat it meansWhat to do
constantGave every project the same score, so their scores say nothing about which is better.every review has the same weighted scoreask them; if they misread the rubric, have them rescore, otherwise top up their projects with another judge
single reviewOnly one review: there’s nothing to compare it with.one review can’t reveal a judge’s leniency or scalegive them more reviews that overlap with other judges
discordantTheir scores don’t rise with everyone else’s on the projects they share.their fitted scale came out at zero or belowread their comments first; more overlapping reviews settle it either way

On the fixture, the page lists twelve judges. It shows names; the fixture ids are here so you can match them to fixtures.json and to the method chapter:

Fixture idName on the pageReviewsFlag
jdg_07Iva Petrova3constant
jdg_01Tomas Varga1single review
jdg_23Anya Sokolova1single review
jdg_04Noor Haddad4discordant
jdg_08Marek Nowak3discordant
jdg_10Hiro Tanaka3discordant
jdg_14Emeka Adeyemi3discordant
jdg_17Bruno Costa2discordant
jdg_18Lars Berg3discordant
jdg_19Mira Kaur4discordant
jdg_21Sana Aziz4discordant
jdg_27Leila Nasser2discordant

jdg_07 gave 4 on every criterion of all three projects. A judge who gives everyone the same score tells you nothing about which project is better, so the model gives them no weight (a per-judge z-score would divide by zero here). The side effect is on the projects they reviewed: one of them, prj_19 (Small Relay), has only two reviews, so its calibrated score rests on the other judge alone.

jdg_01 and jdg_23 wrote one review each. With one review there is nothing to separate a judge’s leniency from the project’s quality, so they get no weight either. jdg_01’s one review is of prj_07, which matters in the first worked example below.

Discordant is not an accusation. It means that on the two to four projects a judge shared with others, their scores went down, or stayed flat, while everyone else’s went up. With so few reviews, one honest disagreement can do that. On the fixture it’s mostly noise, which fits the agreement line. The flags are recomputed on every run, so a judge can move between ok and discordant as reviews come in.

Each judge’s fitted numbers (offset, scale, noise and flag) are in the Judges export, judges.csv.

The ranking

Below the judges is the ranking itself, one row per project that has at least one review in the run. Projects with no reviews at all don’t appear.

ColumnWhat it is
Rankthe calibrated rank among ranked projects; ties, which are rare, go to the higher raw mean
Could bethe 90% bootstrap range for the rank, “a to b”
Raw rankthe rank by plain mean of weighted scores, with “up n” or “down n”
Projectthe title, linked to the project page, and any flag
Reviewssubmitted, complete reviews of this project in the run, flagged judges included
Raw meanthe plain mean of those reviews’ weighted scores, 0 to 1
Calibratedthe calibrated score, mapped back onto the rubric’s 0 to 1 scale

Rank and calibrated score

The calibrated score is the model’s estimate of each project’s quality with every judge’s leniency and scale taken out. It’s put back on the 0 to 1 scale (the average project, plus so many standard deviations of the project means) so it reads like a rubric score, but it’s only comparable within one run. A project seen by few judges, or by noisy ones, is pulled towards the middle rather than trusted on thin evidence.

Could be

This is the column I’d read first. For each run, every project’s reviews are resampled with replacement and the whole model refitted, 200 times, and the column shows the range the project’s rank falls in 90% of the time. A narrow range means the scores place the project consistently; a wide one means they can’t tell it from its neighbours. Where two projects’ ranges overlap, the judges couldn’t really separate them.

On the fixture, most ranges are wide (half of them span 27 places or more), which is the agreement line again, seen project by project.

One caution: the bootstrap can only resample the reviews that exist. A project with two reviews has only three possible resamples, so a narrow range on a low-coverage project is less reassuring than it looks.

Raw rank, up and down

The raw rank orders the same projects by their plain mean. The label next to it compares the two: “up 30” means the project is 30 places higher after calibration than on raw means, “down 4” means 4 places lower, and no label means it didn’t move. A big move with a narrow “could be” is calibration doing its job: a project that drew a harsh judge, say. A big move with a wide “could be” is the model reshuffling noise.

Low coverage

A project with fewer reviews in the run than the event’s k (reviews per project, 3 by default) is flagged low coverage. On the fixture, eight projects from the two unfinished batches have two reviews each:

ProjectTitleReviewersRankCould beRaw rank
prj_10Still Beaconjdg_15, jdg_29165 to 203, down 13
prj_15Copper Orbitjdg_13, jdg_10 (discordant)195 to 3212, down 7
prj_18Open Kilnjdg_24, jdg_04 (discordant)265 to 3619, down 7
prj_19Small Relayjdg_29, jdg_07 (constant)2318 to 2913, down 10
prj_24Glass Beaconjdg_18, jdg_19 (both discordant)left out
prj_29Flat Relayjdg_24, jdg_09183 to 3630, up 12
prj_39Paper Anchorjdg_26, jdg_18 (discordant)281 to 3320, down 8
prj_40Slow Loomjdg_26, jdg_243021 to 3737, up 7

Most of them moved towards the middle. That is the shrinkage working: two reviews, several of them from judges who carry no weight, aren’t enough evidence to keep a project near the top or the bottom. prj_10 had the third-best raw mean on two reviews and ends up 16th, with a range of 5 to 20.

To clear the flag, open Hand out reviews. With k = 3, the plan for the fixture is exactly one top-up for each of the eight. The flag stays until the new reviews are submitted and you run calibration again.

Left out

Some projects are in the table but not ranked. They’re at the bottom, greyed out, with left out: and a reason. Their reviews still help calibrate the judges who wrote them; the project just doesn’t get a rank.

ReasonWhen
duplicatethe project is flagged as a duplicate of an earlier one and nobody has put it back
not submittedthe project has reviews but is a draft
no usable reviewsevery one of its reviews is from a judge who carries no weight

On the fixture there are two:

  • prj_41 Dry Harbour, left out: duplicate. Team CopperLedger submitted Dry Harbour twice, with the same title and repository: prj_07 on 1 March at 04:29 and prj_41 three minutes before the deadline. The import flags the later one. Its four reviews still count towards calibrating their judges. Open Duplicates to confirm it or put it back; if you put it back, run calibration again and it’s ranked like any other project.
  • prj_24 Glass Beacon, left out: no usable reviews. Both of its reviews are from discordant judges (jdg_18 and jdg_19), so the model has nothing to go on. The calibrated score it shows is just the average project, and means nothing. It needs a review from another judge.

Two worked examples

prj_07 Dry Harbour is ranked first, and I wouldn’t announce it. Its raw rank is 31, so calibration moved it up 30 places. It has five reviews, but two are from discordant judges (jdg_19, jdg_21) and one is jdg_01’s single review, so its calibrated score rests on two judges, jdg_12 and jdg_26, who both scored it well relative to how they scored everything else. Its “could be” is 1 to 39: nearly the whole field. On this data, first place means “the two judges who count liked it”, not “it won”.

prj_11 Salt Ledger and prj_34 Iron Switch are the ones I’d trust. They had the best raw means (raw rank 1 and 2) and calibration moved each of them down four places, to 5 and 6. But their ranges are among the narrowest on the page: 1 to 16 and 2 to 12. Whatever the resampling does, the data keeps putting them near the top. If I had to name a shortlist from the fixture, I would build it from the ranges, not the ranks.

Publishing

Until results are published, the results page and the results API answer 404 to everyone except the event’s organizers, who see the latest run with a “Preview” banner.

Publish results freezes the public results to the run on this page:

  • It records the time (on the database’s clock) and the run on the event, and writes a results.publish audit row with the run and its fingerprint.
  • /events/<slug>/results becomes public. It shows each ranked project’s rank, title, team, track, calibrated score, number of reviews, the range it could have landed in, and the fingerprint. Left-out projects aren’t listed. Judge names and flags are never public.
  • /api/events/<slug>/results becomes public too, with the same projects plus each one’s raw rank, raw mean and standard error (see the API).
  • Running calibration again afterwards changes nothing public. The page then offers Publish run n instead, so moving the public result to a newer run is always a deliberate act, and it’s audited.
  • Hide the results again takes them down (results.unpublish in the audit log). Both pages go back to 404 for everyone but organizers.

One thing to watch: the Results and Judges exports always describe the latest run, not the published one. Their run column says which run a file came from.

Before I press publish I check, in order: judges linked, no low-coverage flags left, duplicates decided, the flagged judges looked at, and what the agreement line says. The last one decides how I word the announcement.

Recompute the fingerprint

Anyone with the scores export can check a fingerprint. The export is organizer only, so in practice an organizer downloads it and hands it to whoever wants to check: a team, another organizer, a hackathon judge.

This is what a run hashes, from run_calibration in src/portal/results.py and digest in src/portal/calibration.py:

  • one row per submitted review with every criterion scored, as [review_id, judge_email, project_id, [[criterion_key, value], ...]], the pairs sorted by criterion key;
  • the rows sorted, which puts them in review_id order;
  • serialized with json.dumps(rows, separators=(",", ":")) and hashed with SHA-256.

Every piece of that is a column in scores.csv: review_id, judge_email, project_id (the portal’s numeric id, not the fixture’s prj_ id), and one column per criterion, headed by its key. A review the run skipped has an empty submitted_at or an empty weighted_0_1.

curl -s -H "Authorization: Bearer bb_demo_organizer_5c1e0a" \
  -o scores.csv http://localhost:8080/api/events/sample-hack-2026/export/scores.csv
python3 fingerprint.py scores.csv
# fingerprint.py: recompute a calibration run's input digest from scores.csv
import csv, hashlib, json, sys

FIXED = {"review_id", "judge_email", "project_id", "project", "submitted_at", "weighted_0_1", "comment"}

with open(sys.argv[1], newline="", encoding="utf-8") as f:
    reader = csv.DictReader(f)
    criteria = [c for c in reader.fieldnames if c not in FIXED]
    rows = []
    for r in reader:
        if not r["submitted_at"] or not r["weighted_0_1"]:
            continue  # drafts and incomplete reviews aren't read
        rows.append([int(r["review_id"]), r["judge_email"], int(r["project_id"]),
                     sorted([c, int(r[c])] for c in criteria)])

canon = json.dumps(sorted(rows), separators=(",", ":"))
print(hashlib.sha256(canon.encode()).hexdigest())

On a freshly seeded stack this prints 397dd35f…178f4a, the fixture’s fingerprint. If it doesn’t match a run’s fingerprint:

  • A score changed after the run. That’s what the fingerprint is for. The audit log says which review, when, and who. The fixture’s judging window has no close date, so its judges can still change scores.
  • A judge’s email changed. The export uses the current address.
  • A cell was escaped. The export puts a quote in front of any cell that starts with =, +, - or @, so a spreadsheet won’t run it as a formula. Criterion keys and emails normally never start with those; if one does, strip the leading ' first.

Compare with the fingerprint of the run you care about: the published one is input_digest in the results API, the latest one is on this page and in results.csv.

Architecture

ballotbench is one Django application in front of one Postgres database, started by one docker compose up. Pages are rendered on the server; the JSON API sits next to them and answers the same questions with the same rules. There is no JavaScript framework, no queue beyond a webhook outbox table, no cache and no outside service, because a hackathon portal has hundreds of users, not millions, and every moving part is something a volunteer organizer has to keep running.

browser ──► gunicorn :8080 ──► Django ──► Postgres 18 ◄── deliver_webhooks ──► your receivers
curl    ─┘   (web container)    │         (db container)   (webhooks container, same image)
                                ├─ pages   portal/views, participant, judge, organizer, progress, results,
                                │          oversight, voting, comments, explain, records, embed, webhooks
                                ├─ API     portal/api, judge, results, voting, comments, records, bundles, exports
                                └─ rules   portal/access  +  triggers in the database

Components

ModuleWhat it does
models.pythe schema, with its unique and check constraints
migrations/0002-0005, 0012the triggers (12 in all): deadline, team rules, judging rules, the audit chain, and the voting rules in 0012
access.pyroles, the scoped querysets every view starts from, the status-code rules, the deadline check on the database clock
auth.pybearer tokens (hashed), for the API and, via a middleware, the pages
audit.pywrites one audit row per change, in the change’s transaction
importer.py, commands/seed.pythe idempotent fixture import and the demo accounts
participant.pyteams, invite links, the project form
judge.pythe judge’s queue, the scoresheet, recusal, the judging API
organizer.pyevent settings, tracks, prizes, rubric, role invites
assignment.pythe review planner, a pure function (see JUDGING.md)
progress.pyapplying plans, manual changes, the progress dashboard
calibration.pythe calibration model, bootstrap intervals, the agreement test, a pure module
results.pycalibration runs, publishing, the results pages and API
oversight.pyaudit log page, duplicate resolution, the exports page
exports.pyCSV exports
duplicates.pyduplicate submission detection
voting.py, ratelimit.py, mail.pythe community vote, rate limits counted in Postgres, the offline mail outbox
comments.pyproject comments and their moderation
pairwise.py, explain.pythe Bradley-Terry second opinion; a team’s own “how we were scored” page
records.py, signing.py, embed.py, bundles.py, webhooks.py, tokens.pyT4: signed records and certificates, the embeddable gallery, event bundles, webhooks, personal API tokens

The two pieces with real logic, the planner and the calibration, take plain tuples and return plain objects. They know nothing about Django, so they’re tested on thousands of random inputs in milliseconds, and the database layer around them is thin.

A request, end to end

GET /api/judge/scores?judge=jdg_24 with judge B’s token:

  1. gunicorn hands the request to Django. ATOMIC_REQUESTS opens a transaction for the whole request.
  2. DRF runs BearerTokenAuthentication: SHA-256 the token, look up the hash, refuse a revoked token or an inactive user (401).
  3. api.judge_scores resolves jdg_24 to a user. It isn’t the caller, and the caller organizes no event, so it raises PermissionDenied: 403. A name that matches nobody also gets 403, so the answer doesn’t reveal which judges exist.
  4. Without ?judge=, the queryset starts from access.judge_assignments(user): the caller’s own assignments, in events where they still hold a judge membership. Reviews are filtered from that, never looked up by id.

POST /api/events/sample-hack-2026/projects with the participant’s token:

  1. Authenticated as above.
  2. The caller needs a team in the event, or it’s 403.
  3. access.submissions_closed_reason asks the database for now() against the event’s window. Past the close: 409 with the close time.
  4. If it were open, the write happens inside access.guarded(), a savepoint that turns a trigger’s refusal into a 409 or 422. The deadline trigger checks the same clock again, so a request that races the deadline by a millisecond is still refused.
  5. An audit row is written in the same transaction.

Where each rule is enforced, and why there

RuleApp (for the message)Database (for the guarantee)
deadlinesubmissions_closed_reason, on the DB clockproject_deadline trigger
a judge sees only their own workquerysets start from judge_assignments(user); naming another judge is 403assignments can’t exist for non-judges or own-team projects
one team per person, team sizeform checksunique constraint, team_member_rules trigger with a row lock
no judging your own teamplanner skips itassignment_rules, team_member_zz_not_judge, membership_judge_not_member
scores in rangeform and API validationscore_rules trigger
rubric frozen once scoredpage hides the inputsrubric_frozen trigger
audit log can’t changeno update code existsaudit_readonly, audit_no_truncate, the hash chain
results hidden until published, and while voting is openresults.visible_run, voting.public_tallies; results.publish refuses while voting is open(read rule; no write to guard)
a vote: window, budget, own team, eligible project, confirmed and not voidedvoting.castvote_rules trigger, locking the voter row
one ballot per account and per inboxvoting.normalize_email, request_linkunique constraints on portal_voter
rate limitsratelimit.allow, counted in portal_ratehit(shared across workers because it is in the database)

The database is the last word because a portal has more write paths than anyone remembers: pages, the API, the admin, a management command, a shell at 2 a.m. during the event. The app’s checks are for clear error messages; the triggers are for when someone forgets.

Status codes

The same everywhere, pages and API:

  • 400: the request is malformed: a required field missing, a wrong type, a track from another event.
  • 401: no credentials, or a bad token, on a route that needs them. Pages redirect a browser to the login page instead.
  • 403: you’re signed in but your role can’t do this, or you named another person’s data (?judge=).
  • 404: the object exists but isn’t yours to see: another judge’s assignment, another team’s draft, unpublished results. Saying 403 would confirm it exists.
  • 409: the window for this action is closed, or a conflict of interest.
  • 422: values out of range.

Trade-offs

  • Server-rendered pages, not a SPA. Fewer moving parts, works without JavaScript, one place for the rules. The cost is less interactivity, which a judging form doesn’t need.
  • Triggers in SQL. They’re less familiar to many Django developers than model methods, and they tie us to Postgres. In exchange, the rules hold for every write path, including ones that don’t exist yet. Each trigger’s migration explains it in a docstring.
  • The calibration in pure Python, no numpy. A hackathon has tens of judges and hundreds of reviews: a fit takes 16 ms and 200 bootstrap refits take about 3 seconds. One dependency fewer in the image.
  • No email leaves the box. The portal must run offline, so invite links are shown once to the person who makes them, to send however they like. The one thing that has to be mailed, a voter’s confirmation link, lands in an outbox table that site admins read in the admin. A real deployment sets DJANGO_EMAIL_BACKEND to SMTP.
  • Sessions and tokens side by side. Browsers use sessions (with CSRF); scripts and the checker use bearer tokens (no cookies, so no CSRF risk). Pages accept tokens too so that the isolation probe tests what a browser gets.
  • Two processes. gunicorn with three workers, and the webhooks sender (same image). Nothing else runs in the background: calibration takes a few seconds and runs inside the organizer’s request.

Beyond T2 (tier T4)

ModuleWhat it does
signing.pythe Ed25519 key (one file, made on first use with mode 0600), canonical JSON, sign and verify; a pure module
records.pysigned participation records for judges and participants, /verify, the public key, certificates
embed.pythe frameable gallery /embed/<slug> and /embed.js
bundles.py, commands/export_event, commands/import_eventwhole-event bundles out, and in as a new event through importer.py
webhooks.py, commands/deliver_webhooksthe outbox sender, the address checks, the organizer’s page
  • Records are checkable without us. What’s signed is the record’s canonical JSON (sorted keys, no spaces, UTF-8), so the Python snippet on /verify checks one with nothing but the public key. A record never holds a score: judges’ scores stay private even from the people they judged.
  • Framing is opt-in per view. X_FRAME_OPTIONS = "DENY" stays the default; only the two embed views are exempt, and the embed renders as an anonymous visitor whatever cookies arrive, so it can’t leak a draft into someone else’s page. The iframe reports its height with postMessage, and embed.js accepts it only from that iframe and the portal’s origin.
  • Webhooks are the one background worker. audit.record queues a WebhookDelivery next to the audit row, in the same transaction, so the outbox and the log can’t disagree. A separate webhooks service (same image) sends them: it claims a round, one due delivery per webhook, by leasing them under SKIP LOCKED row locks (so two senders never claim the same one), commits, and sends them all at once with nothing locked. The web process never makes an outbound request.
  • SSRF. A webhook URL must resolve only to public addresses, when it’s saved and again at every send, and the connection goes to the address that was checked (no second DNS lookup to rebind), with one 10 s deadline for the whole attempt and no redirects.

Data model

Postgres 18. Every rule that must hold whatever code path writes a row lives in the database: foreign keys, unique and check constraints, and triggers. The application checks the same rules first so it can answer with a clear 409 or 422, but it doesn’t have to be right for the data to stay right. Models are in src/portal/models.py; triggers are in the migrations 0002 to 0005, plus the voting rules in 0012 (12 triggers in all).

Tables

People and access

portal_user: an account. Signs in with email (unique, and a check constraint keeps it lowercase). is_staff is the global admin role; every other role is per event. email_confirmed_at is set when the person proves they read the inbox (a signed confirmation link, or a completed password reset); account-mode community votes need it.

portal_apitoken: user, label, token_hash (unique), created_at, revoked_at. Only the SHA-256 of a token is stored; the token itself is shown once.

portal_membership: user, event, role (participant, judge or organizer), tracks (many-to-many; for judges, the tracks they may review, empty meaning any), external_id.

  • unique (user, event, role)
  • unique (event, role, external_id): fixture judge ids like jdg_24
  • trigger membership_judge_not_member: nobody becomes a judge of an event they’re on a team in

portal_roleinvite: single-use link that grants judge or organizer in one event. token_hash (unique), expires_at, used_by, used_at, tracks. Check: used_by set only with used_at.

Events

portal_event: slug (unique), name, external_id (unique), submissions_open/close, judging_open/close, voting_open/close, results_published_at, published_run (the calibration run the public results are frozen to), voting_mode (off, account or email), vote_credits (each voter’s quadratic budget), reviews_per_project (k), max_team_size.

  • checks: each window opens before it closes; k ≥ 1; team size ≥ 1

portal_track: event, name, external_id. Unique (event, name) and (event, external_id).

portal_prize: event, optional track, name, description. Unique (event, name).

Teams and projects

portal_team: event, name, external_id. Unique (event, external_id). Names are not unique: the fixture has three different teams called StillTrail.

portal_teammember: team, event, user, joined_at.

  • unique (event, user): one team per person per event
  • trigger team_member_rules: copies the team’s event into event (so the unique constraint means what it says), locks the team row, and refuses a member past max_team_size. The lock makes two simultaneous invite acceptances queue rather than both squeeze in.
  • trigger team_member_zz_not_judge: a judge of the event can’t join a team in it

portal_teaminvite: team, token_hash (unique), created_by, expires_at, used_by, used_at. Single use, 72 hours or until submissions close, whichever is first.

portal_project: event, team, track, title, tagline, summary, description, repo_url, demo_url, tags (text array), status (draft or submitted), submitted_at, duplicate_of (self), duplicate_cleared (an organizer said it isn’t one; never flagged again), external_id.

  • unique (event, external_id)
  • check: submitted ⇔ submitted_at set
  • check: not its own duplicate
  • index (event, status) for the gallery
  • trigger project_same_event: team and track belong to the project’s event
  • trigger project_deadline: see below

Judging

portal_rubriccriterion: event, key, name, weight, min_value, max_value, position.

  • unique (event, key); checks: weight > 0, min < max
  • trigger rubric_frozen: once any score exists in the event, no criterion can be added, deleted, reweighted or re-ranged. Renaming is allowed.

portal_judgeassignment: event, judge, project, status (pending, done, recused), source (seed, auto, manual).

  • unique (judge, project)
  • trigger assignment_rules: the project is in the event, the judge has a judge membership in it, and the judge isn’t on the project’s team

portal_review: one per assignment (assignment unique). comment, submitted_at (null while a draft).

portal_score: review, criterion, value.

  • unique (review, criterion)
  • trigger score_rules: the criterion belongs to the review’s event and the value is inside its range
  • criterion is ON DELETE RESTRICT: a scored criterion can’t be deleted alone, but an event can be deleted whole

Calibration

portal_calibrationrun: event, created_at, created_by, method, params (JSON: priors and the rubric weights used), input_digest (SHA-256 of exactly the scores read), connected, converged, signal_p, n_reviews.

portal_calibratedproject: run, project, n_reviews, raw_mean, quality (q), display (q on the 0..1 scale), se, rank, raw_rank, rank_low, rank_high (90% bootstrap interval), excluded (why it isn’t ranked: duplicate, not submitted, no usable reviews). Unique (run, project).

portal_judgecalibration: run, judge, n_reviews, offset, scale, noise, flag (ok, constant, single_review, discordant). Unique (run, judge).

A run is never updated. Publishing points event.published_run at one.

Community voting and comments

portal_voter: one ballot in one event. event, user (account mode) or email as typed plus email_normalized (email mode), token_hash (the SHA-256 of the one-time confirmation link), confirmed_at, ip, user_agent, voided_at, voided_reason.

  • unique (event, user) and (event, email_normalized): one ballot per account and per inbox
  • checks: an account or an address; a voided ballot has a reason

portal_vote: voter, project, votes (≥ 1; a zero is no row). Unique (voter, project).

  • trigger vote_rules (migration 0012): locks the voter row, then refuses the write outside the voting window (BB409), for a voided or unconfirmed voter (BB409), for a project that isn’t a submitted, non-duplicate project of the voter’s event (BB422), for the voter’s own team (BB423), or when the sum of votes² would pass vote_credits (BB409).

portal_comment: project, author, body, created_at, hidden_at, hidden_by.

  • checks: body not empty and at most 2000 characters

portal_ratehit: key, created_at. One counted action for a rate limit; ratelimit.allow() deletes a key’s rows once they leave its window.

portal_outboundemail: to, subject, body, created_at. Mail the portal would send, kept for site admins to read (portal.mail.OutboxBackend).

The audit log

portal_auditlog: seq, ts, actor, event, action, object_type, object_id, before and after (JSON), ip, prev_hash, row_hash.

  • trigger audit_append (BEFORE INSERT): takes a transaction-scoped advisory lock, sets seq to the next number, ts to the clock, prev_hash to the previous row’s hash, and row_hash = sha256(prev_hash | seq | ts | actor | event | action | object | before | after | ip).
  • triggers audit_readonly and audit_no_truncate: UPDATE, DELETE and TRUNCATE are refused.
  • actor and event have no foreign-key constraint on purpose: deleting a user or an event must not delete, or be blocked by, its history.
  • manage.py verify_audit recomputes the chain and names the first row that doesn’t fit.

Why seq and not the primary key: ids are handed out by a sequence before the trigger takes its lock, so two concurrent inserts can commit in the opposite order to their ids. seq is assigned under the lock.

The deadline, in the database

project_deadline runs before every INSERT, UPDATE and DELETE on portal_project. If the database clock (now()) is past the event’s submissions_close, it refuses a new project, a deleted one, or any change to content (title, text, links, tags, track, team, status). It allows changes to duplicate_of and duplicate_cleared, which are organizer bookkeeping.

The one bypass is the session setting ballotbench.import, which only the fixture importer and delete_event set, with SET LOCAL so it ends with their transaction, and both write an audit row saying so.

Refusals and what the API returns

Triggers raise with their own SQLSTATE, which access.guarded() maps to a status code:

SQLSTATEMeaningHTTP
BB409a window is closed (deadline, frozen rubric)409
BB410team is full409
BB423conflict of interest409
BB422an inconsistent reference (wrong event, out-of-range score)422
BB403the audit log is append only403

Getting data in

  • The fixture: manage.py seed on every boot. Imported rows keep their fixture ids in external_id, with UNIQUE (event, external_id), so a second run inserts nothing and never overwrites what people changed. Fixture projects come in as submitted, through the import bypass, because they were submitted before a close date that has passed.
  • Another event in the fixture’s shape: portal.importer.import_event(data) is the same importer; manage.py seed --fixtures path.json loads one.

Getting data out

  • CSV, organizer only, audited, at every stage: registrations, teams, projects, assignments, scores (each criterion and the weighted score), results (raw and calibrated, rank intervals, the run’s digest), judges (per-judge calibration), audit (with the hash chain). GET /api/events/<slug>/export/<kind>.csv.
  • JSON: everything the pages show is in the API (/api/docs).
  • The database itself: plain Postgres, pg_dump works.

Beyond T2 (tier T4)

portal_webhook: event, url, secret (64 hex characters; it keys the HMAC, so it’s stored as is, shown once and never written to the audit log), actions (text array of action prefixes, empty meaning every change), active, created_by (who it sends on behalf of: it pauses once they no longer organize the event, and whoever resumes it takes it over), created_at.

portal_webhookdelivery: the outbox. webhook, action, payload (JSON: the audit row’s action, event, actor, object, before and after), status (pending, delivered, failed), attempts, next_attempt_at (defaults to the database clock; while a sender has it claimed, the end of its lease), last_error, created_at. Index (status, next_attempt_at) for the sender. A row is written by audit.record in the change’s transaction, so it rolls back with it. Failed sends wait 30 s, then twice as long each time; the eighth failure is final.

Signed records aren’t stored: each is built from the tables above when asked for, and the issuance is an audit row (record.issue). The signing key is a file (BALLOTBENCH_SIGNING_KEY_FILE, /data/signing_key.pem in Docker), not a table, so a database dump doesn’t carry it.

Event bundles (bundles.py) are the fixture’s shape plus format, prizes, criteria (with weights and ranges), assignments (those without a review) and, per row, everything the fixture leaves out: event dates, judges’ tracks, every project field with status and duplicate_of, and each review’s submitted flag, time, status and comment. Rows are named by external_id, or prj_<pk> and the like. Importing one as a new event goes through import_event(..., new=True) with the deadline bypass, in one transaction; URLs, emails and every enum are validated first, and the bypass is switched off again when the import ends rather than when the request does.

Threat model

Who attacks a hackathon portal, how, and what ballotbench does about it. The stakes are small prizes and bragging rights, so most attackers are participants with a browser, a few friends and an evening, not nation states. Each entry says what stops the attack, where that lives in the code, and what doesn’t stop it.

People

  • Participants want their project to win, want to see others’ work early, and want more time.
  • Their friends will vote, and will make a few extra accounts if asked.
  • Judges are mostly honest; a few favour a friend or want to see how their scores compare.
  • Organizers are trusted with the event, but not with each other’s secrets or with rewriting history quietly.
  • Anyone on the internet can reach the public pages and the API.

Attacks

Sybil voters: one person, many ballots

  • Stops it: one ballot per account and per inbox, as unique constraints on portal_voter. Email addresses are normalized first (voting.normalize_email: lowercase, +tag dropped, Gmail dots and googlemail.com folded), so a.b+1@googlemail.com is ab@gmail.com. A second spelling of an inbox gets no second ballot and is logged as vote.duplicate_refused; its link goes to the spelling just typed, and replaces the last one, so whoever typed alice+x@ first doesn’t get alice@’s links. An address with a quoted local part ("a.b"@gmail.com) is refused rather than normalized. Accounts must confirm their address by a link before they can vote. Sign-ups, link requests and resets are rate limited per address (the written limits times BALLOTBENCH_ADDRESS_LIMIT_SCALE, 10 by default, so a venue behind one NAT address isn’t locked out) and link requests per inbox (3 an hour). The organizer’s abuse panel (voting.abuse_report) flags networks (/24, or /64 for IPv6) with three or more voters, accounts created less than an hour before their first ballot, and three or more identical ballots. Organizers void a ballot with a reason; it leaves the tallies for good and the audit log says who did it and why.
  • Doesn’t stop it: someone with many real inboxes, a catch-all domain, or a botnet of addresses. There is no CAPTCHA, phone check or proof of personhood. Flags are never acted on automatically, because an office or a campus shares one network and friends vote alike; a person decides. Quadratic voting limits how much any one ballot can do, not how many there are.

Ballot stuffing: one ballot, more weight

  • Stops it: the vote_rules trigger (migration 0012) locks the voter’s row and refuses any write that would take the sum of votes² past vote_credits, so two tabs or a replayed request can’t overspend. It also refuses votes outside the window on the database clock, votes for a project that isn’t a submitted, non-duplicate project of the voter’s event (so no voting across events), votes for your own team, and votes by unconfirmed or voided voters. The app (voting.cast) checks the same things first so it can say why; it never trusts a cost sent by the client. Ballot writes are rate limited per address and per voter.
  • Doesn’t stop it: an email voter on a team whose members signed up with a different address. Own-team votes are refused for accounts in the database, and for email voters only when a team member’s address reaches the same inbox, which only the app can check.

Reading tallies or results early

  • Stops it: tallies are shown only to the event’s organizers until results are published (voting.public_tallies), results.publish refuses while voting is open, and published results are hidden again if voting reopens (results.visible_run), so nobody votes with the judges’ ranking in front of them. Unpublished results are a 404, not a 403. While voting is open organizers see how many ballots there are and what each voter spent, not the tallies, and a void can’t be undone: with live tallies and an undo, voiding one ballot, reading the tallies and counting it again would show what that voter chose.
  • Doesn’t stop it: an organizer telling people once voting closes. An organizer who voids a ballot after the close can see from the tallies what it held, at the price of throwing it away; one who moves the close date to read the tallies and then reopens voting can compare two reads. Both leave dated rows in the audit log.

Pushing a rival out as a duplicate

  • Stops it: a flagged duplicate leaves judging and the ballot, so the detector (duplicates.find_duplicates) only compares submitted projects and calls the one submitted later the copy, by submitted_at, which only the server sets. A submit only ever flags the project being submitted (duplicates.flag_on_submit), so copying another team’s public title and repo into an old draft flags the copy, not the original. An organizer’s “it’s a different project” sets duplicate_cleared, and the detector never flags that project again.
  • Doesn’t stop it: a team that submits a placeholder early and edits it into a copy afterwards looks earlier to a whole-event scan. Submits never run one; only an organizer’s import does, and its flags are on the duplicates page for a person to judge.

Scraping drafts

  • Stops it: every project query goes through access.visible_projects: drafts are visible to their own team and the event’s organizers only, and anyone else gets a 404 on the page and the API. The gallery lists submitted projects only.
  • Doesn’t stop it: a team member sharing their screen, or a public repository linked from the draft.

Peeking at peer scores

  • Stops it: judge queries start from access.judge_assignments(user), so another judge’s review is a 404; naming another judge in /api/judge/scores?judge= is a 403, never an empty list. Review comments go to organizers only. scripts/isolation_curl.sh tries all of this over HTTP.
  • Doesn’t stop it: judges talking to each other.

Judge collusion

  • Stops it: conflict-of-interest triggers (team_member_zz_not_judge, membership_judge_not_member, assignment_rules) keep a judge off their own team’s project and out of any team in an event they judge. Each project gets k reviews from different judges, and every review is in the audit log.
  • Doesn’t stop it: two judges who agree to push a project look like two judges who agree. Calibration can’t tell them apart (JUDGING.md).

Deadline gaming

  • Stops it: the project_deadline trigger (migration 0002) refuses creating, deleting or editing a project after submissions_close, on the database’s clock, whatever path the write takes: page, API, admin or shell. The client’s clock is never asked. Moving the deadline is an audited event.update.
  • Doesn’t stop it: pushing to the linked repository after the deadline. The portal records submitted_at; checking the repository’s history against it is up to the judges.

Tampering with results or the audit log

  • Stops it: published results are frozen to one calibration run, and each run stores the SHA-256 of the exact scores it read (input_digest). The audit log is append only (audit_readonly, audit_no_truncate) and hash chained; manage.py verify_audit names the first row that doesn’t fit. Admin writes are audited like any other.
  • Doesn’t stop it: someone with the database superuser password can disable the triggers and rewrite the whole chain from any point onwards. Keep a copy of the latest row_hash somewhere else (the audit export has it) and a rewrite shows.

CSV formula injection

  • Stops it: every exported cell that a spreadsheet would read as a formula (starting with =, +, -, @, tab or carriage return) and isn’t a plain number is prefixed with ' (exports.cell), so a project called =HYPERLINK(...) opens as text.
  • Doesn’t stop it: a spreadsheet told to ignore that, or someone pasting cells into a formula by hand.

Comments as a weapon

  • Stops it: bodies are plain text, escaped by Django’s autoescape (no |safe anywhere), at most 2000 characters (checked by the serializer and by a database constraint), and rate limited to 10 per account per 10 minutes. The event’s organizers hide a comment and it disappears for everyone else; the row and the audit trail stay.
  • Doesn’t stop it: abuse that’s within the rules until an organizer reads it. There is no filter or pre-moderation.
  • Stops it: team and role invites are single use, only their hash is stored, team invites expire in 72 hours (or at the deadline) and role invites in 7 days, and both can be revoked. Voting links are single use, stored as a hash, and replaced by the next request. A voting link opened with GET only shows a button, so a mail scanner that fetches it doesn’t use it up.
  • Doesn’t stop it: whoever gets a leaked link first. A leaked organizer invite is a new organizer; check the Organizers list on the Settings page and the role.accept rows in the audit log.

Token and session theft

  • Stops it: API tokens are random, shown once, stored as SHA-256, and revocable by their owner at /me/tokens (or by an admin). Session cookies are HttpOnly and SameSite=Lax, every session POST needs a CSRF token, and no page can be framed except the read-only embed (below). Logins are limited per address (20 per person per 10 minutes, times BALLOTBENCH_ADDRESS_LIMIT_SCALE for a venue behind one NAT) and to 10 attempts per account per 10 minutes from anywhere, so a password can’t be guessed at network speed.
  • Doesn’t stop it: tokens don’t expire until revoked. Plain HTTP is the default so the laptop demo works; behind TLS, DJANGO_SECURE=1 turns on secure cookies, HSTS and the redirect to HTTPS. A botnet still gets 10 guesses per account every 10 minutes; a strong password makes that useless.

Flooding

  • Stops it: the rate limits above, counted in Postgres (ratelimit.allow) so every worker shares them. IPv6 is limited per /64, because one subscriber usually holds a whole /64.
  • Doesn’t stop it: a real denial of service. Put a proxy in front, and set BALLOTBENCH_TRUSTED_PROXIES to the number of proxies so the limits and the audit log see the visitor’s address (read from X-Forwarded-For, counting from the right, so a client can’t choose it), not the proxy’s.

Webhooks as a way into the private network

  • Stops it: a webhook URL must be http(s) and every address its name resolves to must be public (webhooks.public: is_global, not multicast, not site-local fec0::/10; an IPv6 address carrying an IPv4 one, mapped, compatible, translated or NAT64 64:ff9b::/96, is judged by the IPv4 one). That’s checked when the organizer saves it and again right before each send, and the connection goes to the address that was checked, so a DNS answer that changes between the check and the request (DNS rebinding) can’t redirect it. Redirects aren’t followed, and each delivery is signed with HMAC-SHA256 so the receiver can tell it’s from this portal.
  • Doesn’t stop it: an organizer choosing to deliver to their own network with BALLOTBENCH_WEBHOOKS_ALLOW_PRIVATE=1. Webhook payloads carry what the audit log carries for that event, so a webhook URL is a copy of the log: only organizers can add one, and adding one is itself audited. It keeps sending only while the organizer who added it (or who last resumed it) still organizes the event or is staff; otherwise the next change pauses it, with a webhook.pause audit row saying why.

A webhook receiver holding up the sender

  • Stops it: one deadline of 10 seconds per attempt, for connecting, sending and reading the answer together (a socket timeout alone counts each read afresh, so a receiver sending a byte every few seconds could hold the sender for ever). Only the status line is read, at most 16 KiB looking for it, and never the body. Sends happen with no transaction or row lock held: the sender claims a round of due deliveries, one per webhook, by moving each next_attempt_at a lease (2 minutes) ahead under SKIP LOCKED, commits, sends them all at once, then records each result. A sender that dies mid-send leaves its deliveries due again when the lease runs out.
  • Doesn’t stop it: a slow receiver still delays the other webhooks’ deliveries by up to one deadline per round, since a round waits for its slowest send. The DNS lookup before each send isn’t under the deadline; the system resolver’s own timeouts bound it.

The embed as a window in

  • Stops it: /embed/<slug> is the only page any site may frame (frame-ancestors *, no X-Frame-Options); everything else is DENY. It renders as an anonymous visitor whatever cookie or token comes with the request, so it never shows a draft or unpublished results, and it has no forms, so framing it can’t trick anyone into clicking something that changes state.
  • Doesn’t stop it: someone embedding a public gallery where you’d rather they didn’t. It’s public anyway.

Forged records and certificates

  • Stops it: records are Ed25519 signatures over canonical JSON; the public key is at /.well-known/ballotbench-signing-key, and /verify checks one without needing an account. Changing one character of a record makes it fail. Records never contain scores. The private key lives on the data volume, created with mode 0600, never in the repository.
  • Doesn’t stop it: someone who can read the data volume can sign anything. There’s no revocation list and no key rotation yet: a new key makes every old record fail to verify.

The API

The JSON API sits next to the pages and answers the same questions with the same rules: every lookup goes through the same scoped querysets in portal/access.py, and every write goes through the same database triggers. Anything a page can tell you, the API can too, with one exception: the organizer’s tools (settings, invitations, handing out reviews, calibration, duplicates) are pages only.

Every example below is a real request against the demo stack on http://localhost:8080, with the demo tokens from .dogfood.toml. The read-only ones you can paste as they are. The ones that write change the demo data, so run them on a stack you’re happy to reset with docker compose down -v.

Authentication

Bearer tokens

Scripts authenticate with a token in the Authorization header:

curl -H "Authorization: Bearer bb_demo_judge_a_8d24f1" http://localhost:8080/api/judge/scores

A token is bb_ followed by 43 random characters. The portal stores only its SHA-256, so a token is shown once, when it’s made, and a leaked database doesn’t leak working tokens. A token acts as its user, with all of that user’s roles; there are no scopes. A revoked token, or one whose user has been deactivated, is refused.

The demo seed creates four tokens with fixed values, so the checker and this book can use them. They’re public, which is why a real deployment turns them off (see Running it for real).

WhoTokenRoles
organizerbb_demo_organizer_5c1e0aorganizer of both seeded events
judge_abb_demo_judge_a_8d24f1jdg_24 in the fixture event, judge of the open demo event
judge_bbb_demo_judge_b_3a9e77jdg_29 in the fixture event
participantbb_demo_participant_61b0c4team NorthKiln in the fixture event, participant in the demo event

judge_a and judge_b share no project, so each has scores the other must not see. Anyone signed in can issue and revoke their own tokens at /me/tokens (linked from My events); a token is shown once and only its hash is stored.

Browser sessions

A browser that’s signed in can call the API with its session cookie. Then Django’s CSRF protection applies: a POST or PATCH needs the X-CSRFToken header with the value of the csrftoken cookie, or it’s refused with 403. Bearer requests need no CSRF token, because a browser never attaches an Authorization header on its own, so a forged cross-site request can’t carry one.

Pages accept tokens too

The HTML pages accept the same bearer tokens, for reading only. A GET with a token sees exactly what that user would see in a browser, and no session is created; this is what lets the isolation probe test the pages with curl. Anything that would change something through a page (a POST) is refused with 403 when it comes with a token, and so are /me/tokens and the admin: a leaked token can’t mint more tokens. Scripts that write use the API. A bad token on a page gets a plain-text 401.

Status codes

The same contract holds on the pages and in the API.

CodeWhenExample
200, 201it worked; 201 when a project was created
400the request body is malformed: a missing required field, a wrong type, a track from another event{"title": ["This field is required."]}
401no credentials, or a bad token, on a route that needs them (API responses carry WWW-Authenticate: Bearer; pages send a browser to the login page instead){"detail": "invalid or revoked token"}
403you’re signed in but your role can’t do this, or you named someone else’s data{"detail": "judges can only read their own scores"}
404the thing doesn’t exist, or it exists but isn’t yours to see: another judge’s assignment, another team’s draft, unpublished results{"detail": "Not found."}
409the window for this is closed, or a conflict of interest{"detail": "Submissions for Sample Hack 2026 closed at 2026-03-01 18:00 UTC."}
422the values are wrong: a score out of range, an unknown criterion, submitting a review with a criterion missing{"detail": "Quality must be between 1 and 5"}

The line between 403 and 404 is deliberate. If you may know a thing exists but can’t touch it (another team’s submitted project, which is in the public gallery), you get 403. If you may not even know it exists (another judge’s assignment, another team’s draft), you get 404, since a 403 would confirm it’s there.

A 409 or 422 can come from the app’s own check or from a database trigger. The app checks first so it can give a clear message; the trigger is the backstop, and access.guarded() turns its refusal into the same status code. Errors are always JSON: {"detail": "..."}, or a field-by-field object for a 400.

Endpoints

MethodPathWho may call it
GET/api/eventsanyone
GET/api/events/<slug>anyone
GET/api/events/<slug>/projectsanyone; drafts only for their team and organizers
POST/api/events/<slug>/projectsa member of a team in the event
GET/api/projects/<id>anyone who can see the project
PATCH/api/projects/<id>the project’s team
GET/api/judge/assignmentsjudges
POST/api/judge/assignments/<id>/reviewthe judge the assignment belongs to
GET/api/judge/scoresjudges, for themselves; organizers, for their events’ judges
GET/api/events/<slug>/resultsanyone once published; before that, the event’s organizers
GET/api/events/<slug>/export/<kind>.csvthe event’s organizers
GET/api/events/<slug>/results/signedanyone once published: the ranking signed with the portal’s Ed25519 key
GET, POST/api/events/<slug>/ballotvoters, in events that vote by signed-in account (confirmed address)
GET, POST/api/projects/<id>/commentsanyone reads; signed-in users post
GET/api/judge/record?event=<slug>a judge, for their own signed record; organizers may name a judge of their event
GET/api/participant/record?event=<slug>a participant, for their own signed record
POST/api/verifyanyone: checks a signed record or signed results
GET/.well-known/ballotbench-signing-keyanyone: the public key records and results are signed with
GET/api/events/<slug>/export/bundle.jsonthe event’s organizers: the whole event as one file
POST/api/events/import?slug=<new>site admins: a bundle back in as a new event
GET/api/schemaanyone
GET/api/docsanyone

Site admins (is_staff) count as organizers of every event.

Events

List events

GET /api/events: every event, the latest deadline first, with its windows, tracks and rubric. No authentication.

curl http://localhost:8080/api/events
[
  {
    "slug": "demo-open",
    "name": "Demo Hack (open)",
    "...": "..."
  },
  {
    "slug": "sample-hack-2026",
    "name": "Sample Hack 2026",
    "description": "",
    "submissions_open": "2026-02-26T00:00:00Z",
    "submissions_close": "2026-03-01T18:00:00Z",
    "judging_open": "2026-03-01T18:00:00Z",
    "judging_close": null,
    "voting_open": null,
    "voting_close": null,
    "results_published_at": null,
    "reviews_per_project": 3,
    "max_team_size": 4,
    "tracks": [{"id": 1, "name": "Developer tools"}, {"id": 2, "name": "Data and analytics"}, "..."],
    "criteria": [
      {"key": "functionality", "name": "Functionality", "weight": "1.000", "min_value": 1, "max_value": 5},
      {"key": "quality", "name": "Quality", "weight": "1.000", "min_value": 1, "max_value": 5},
      {"key": "innovation", "name": "Innovation", "weight": "1.000", "min_value": 1, "max_value": 5}
    ]
  }
]

All times are UTC. A null close means the window never closes. Weights are decimals, sent as strings so they don’t lose precision.

One event

GET /api/events/<slug>: the same object for one event, or 404.

curl http://localhost:8080/api/events/sample-hack-2026

Projects

A project looks like this in every response:

{
  "id": 1,
  "event": "sample-hack-2026",
  "team": "NorthKiln",
  "track": 4,
  "title": "Glass Signal",
  "tagline": "",
  "summary": "One line of what it does.",
  "description": "",
  "repo_url": "https://example.org/repo/01",
  "demo_url": "",
  "tags": [],
  "status": "submitted",
  "submitted_at": "2026-02-27T04:08:00Z",
  "duplicate_of": null
}

track is a track id from the event. status is draft or submitted. duplicate_of is the id of the earlier project this one repeats, if it’s been flagged; flagged projects stay listed but aren’t ranked.

List an event’s projects

GET /api/events/<slug>/projects: the event’s submitted projects, in id order, for anyone. Signed in, you also see your own team’s drafts, and an organizer sees every draft in the event.

curl http://localhost:8080/api/events/sample-hack-2026/projects

On the fixture that’s 41 projects, including the duplicate prj_41 (id 41, "duplicate_of": 7).

One project

GET /api/projects/<id>: one project, if you can see it. Another team’s draft is a 404, not a 403.

curl http://localhost:8080/api/projects/1

Create a project

POST /api/events/<slug>/projects: creates a project for your team in that event. You need to be on a team in the event; teams are made on the event’s page, since the API has no team routes.

FieldTypeNotes
titlestringrequired
tagline, summary, descriptionstringoptional
repo_url, demo_urlURLoptional
tagslist of stringstrimmed, lowercased and de-duplicated; at most 10
tracktrack idmust be a track of this event
submitbooleantrue submits it; otherwise it’s saved as a draft
OutcomeCode
created201, with the project
not signed in401
no team in this event403 you need a team in this event to submit a project
before the window opens or after it closes409, with the time
a bad field400

Each call makes a new project; a team can hold more than one (the fixture’s team CopperLedger has two). When a project is submitted, the portal checks it against the event’s earlier projects and flags it if it repeats one.

This is the request the acceptance checker makes. The fixture event closed on 1 March 2026, so it’s refused:

curl -X POST http://localhost:8080/api/events/sample-hack-2026/projects \
  -H "Authorization: Bearer bb_demo_participant_61b0c4" \
  -H "Content-Type: application/json" \
  -d '{"title": "late", "summary": "x"}'
{"detail": "Submissions for Sample Hack 2026 closed at 2026-03-01 18:00 UTC."}

with status 409. On the open demo event, once the participant has started a team on the event’s page, the same call creates a draft:

curl -X POST http://localhost:8080/api/events/demo-open/projects \
  -H "Authorization: Bearer bb_demo_participant_61b0c4" \
  -H "Content-Type: application/json" \
  -d '{"title": "Night Owl Radio", "tagline": "Offline-first radio for field teams",
       "summary": "Mesh radio notes.", "repo_url": "https://example.org/night-owl",
       "tags": ["radio", "mesh"]}'
{
  "id": 42,
  "event": "demo-open",
  "team": "Night Owls",
  "track": null,
  "title": "Night Owl Radio",
  "tagline": "Offline-first radio for field teams",
  "summary": "Mesh radio notes.",
  "description": "",
  "repo_url": "https://example.org/night-owl",
  "demo_url": "",
  "tags": ["radio", "mesh"],
  "status": "draft",
  "submitted_at": null,
  "duplicate_of": null
}

Edit or submit a project

PATCH /api/projects/<id>: changes the fields you send and leaves the rest. Only members of the project’s team may; send "submit": true to submit it. submitted_at is set from the database’s clock, and there’s no way to unsubmit through the API.

OutcomeCode
saved200, with the project
you can’t see the project404
not signed in401
you can see it but it isn’t your team’s403 only the project's team can edit it
the submission window is closed409
a bad field400
curl -X PATCH http://localhost:8080/api/projects/42 \
  -H "Authorization: Bearer bb_demo_participant_61b0c4" \
  -H "Content-Type: application/json" \
  -d '{"demo_url": "https://example.org/night-owl/demo", "submit": true}'
{"id": 42, "...": "...", "demo_url": "https://example.org/night-owl/demo",
 "status": "submitted", "submitted_at": "2026-09-28T05:13:51.891600Z", "duplicate_of": null}

After the deadline, editing your own project is a 409, and editing another team’s is a 403 (prj_02 is another team’s):

curl -X PATCH http://localhost:8080/api/projects/1 \
  -H "Authorization: Bearer bb_demo_participant_61b0c4" \
  -H "Content-Type: application/json" -d '{"title": "late"}'
# 409 {"detail": "Submissions for Sample Hack 2026 closed at 2026-03-01 18:00 UTC."}

curl -X PATCH http://localhost:8080/api/projects/2 \
  -H "Authorization: Bearer bb_demo_participant_61b0c4" \
  -H "Content-Type: application/json" -d '{"title": "mine"}'
# 403 {"detail": "only the project's team can edit it"}

Judging

Your assignments

GET /api/judge/assignments: your own assignments in every event you judge, and nobody else’s. 403 judges only if you judge no event.

curl -H "Authorization: Bearer bb_demo_judge_a_8d24f1" http://localhost:8080/api/judge/assignments
[
  {"id": 16, "event": "sample-hack-2026", "project": 6, "project_title": "Dry Compass", "status": "done"},
  {"id": 38, "event": "sample-hack-2026", "project": 12, "project_title": "Open Beacon", "status": "done"},
  "..."
]

status is pending, done or recused. Stepping aside from a project (recusal) is on the scoresheet page, not in the API.

Score an assignment

POST /api/judge/assignments/<id>/review: saves your scores for one of your own assignments.

FieldTypeNotes
scoresobjectcriterion key to an integer in that criterion’s range
commentstringoptional, up to 5000 characters; organizers see it, other judges never do
submitbooleantrue submits; otherwise it’s a draft. Submitting needs every criterion.
OutcomeCode
saved200
not your assignment, whatever the id404
judging hasn’t opened, or has closed409, with the time
you stepped aside from this project409
a score out of range, an unknown criterion, or submitting with one missing422
a score that isn’t an integer400

You can save and resubmit as often as you like while judging is open. Every save is in the audit log with the scores before and after.

curl -X POST http://localhost:8080/api/judge/assignments/16/review \
  -H "Authorization: Bearer bb_demo_judge_a_8d24f1" \
  -H "Content-Type: application/json" \
  -d '{"scores": {"functionality": 3, "quality": 4, "innovation": 5}, "comment": "Clear demo.", "submit": true}'
{"assignment": 16, "submitted_at": "2026-09-28T05:13:51.974376Z", "weighted_total": 0.75}

weighted_total is the review’s score on the 0 to 1 scale (see the method). The fixture event’s judging window has no close, so this really does change the fixture’s scores, and with them the calibration fingerprint.

The same assignment, as another judge, doesn’t exist:

curl -X POST http://localhost:8080/api/judge/assignments/16/review \
  -H "Authorization: Bearer bb_demo_judge_b_3a9e77" \
  -H "Content-Type: application/json" -d '{"scores": {"quality": 5}}'
# 404 {"detail": "No JudgeAssignment matches the given query."}

Your scores

GET /api/judge/scores: your own reviews, drafts included, with every criterion’s score and the weighted total.

curl -H "Authorization: Bearer bb_demo_judge_a_8d24f1" http://localhost:8080/api/judge/scores
[
  {
    "assignment": 16,
    "event": "sample-hack-2026",
    "project": 6,
    "project_title": "Dry Compass",
    "judge": "diego.herrera@example.org",
    "status": "done",
    "submitted_at": "2026-03-01T18:00:00Z",
    "comment": "Runs clean.",
    "scores": {"functionality": 2, "quality": 3, "innovation": 5},
    "weighted_total": 0.5833333333333334
  },
  "..."
]

Two query parameters:

  • event=<slug> keeps one event’s reviews.
  • judge=<judge> names a judge, by fixture id (jdg_24), email or user id. A judge may name only themselves. Naming anyone else is a 403, never an empty list, so a refusal can’t be mistaken for “no scores”; and a name that matches nobody is a 403 too, so the answer doesn’t reveal which judges exist. An organizer may name any judge and gets that judge’s reviews in the events they organize (404 no such judge if the name matches nobody).
CallerResult
a judge, no judge=200, their own reviews
a judge naming another judge, or nobody403 judges can only read their own scores
someone who judges nothing, no judge=403 judges only
an organizer naming a judge200, that judge’s reviews in the organizer’s events
no credentials401

The acceptance checker’s peer probe:

curl -H "Authorization: Bearer bb_demo_judge_b_3a9e77" "http://localhost:8080/api/judge/scores?judge=jdg_24"
# 403 {"detail": "judges can only read their own scores"}

curl -H "Authorization: Bearer bb_demo_organizer_5c1e0a" \
  "http://localhost:8080/api/judge/scores?judge=jdg_24&event=sample-hack-2026"
# 200, jdg_24's eleven fixture reviews

Results

GET /api/events/<slug>/results: the ranked results. Until they’re published this is a 404 for everyone except the event’s organizers, who get the latest calibration run (with "published_at": null). After publishing, everyone gets the published run, and later runs change nothing here.

curl http://localhost:8080/api/events/sample-hack-2026/results
# 404 {"detail": "Not found."} until an organizer publishes

curl -H "Authorization: Bearer bb_demo_organizer_5c1e0a" http://localhost:8080/api/events/sample-hack-2026/results
{
  "event": "sample-hack-2026",
  "published_at": null,
  "run": 1,
  "input_digest": "397dd35f376eb052b0dfcd8fda8309f46bad4771e9caa4e78035175807178f4a",
  "method": "offset-scale-noise/v1",
  "projects": [
    {
      "rank": 1,
      "raw_rank": 31,
      "project": 7,
      "title": "Dry Harbour",
      "team": "CopperLedger",
      "calibrated": 0.8321,
      "se": 0.0304,
      "raw_mean": 0.5833,
      "reviews": 5,
      "rank_interval": [1, 39]
    },
    "..."
  ]
}

That’s the fixture after one calibration run and before publishing. Only ranked projects are listed; duplicates and projects with no usable reviews are left out. rank_interval is the 90% bootstrap range the page shows as “could be”; se is the model’s own standard error, which is narrower because it treats every judge’s habits as known. input_digest is the fingerprint of the scores the run read (how to check it).

CSV exports

GET /api/events/<slug>/export/<kind>.csv: one export per stage of the event, for its organizers. 401 without credentials, 403 for anyone else, 404 for a kind that doesn’t exist. Every download writes an export.csv row to the audit log.

curl -H "Authorization: Bearer bb_demo_organizer_5c1e0a" \
  http://localhost:8080/api/events/sample-hack-2026/export/scores.csv | head -3
review_id,judge_email,project_id,project,submitted_at,functionality,quality,innovation,weighted_0_1,comment
1,marek.nowak@example.org,1,Glass Signal,2026-03-01T18:00:00+00:00,2,4,2,0.4167,Runs clean.
3,pavel.ivanov@example.org,1,Glass Signal,2026-03-01T18:00:00+00:00,2,5,3,0.5833,Solid.
KindOne row perColumns
registrationsmembershipuser_email, name, role, tracks
teamsteam memberteam_id, external_id, team, member_email, joined_at
projectsproject, drafts includedproject_id, external_id, title, team, track, status, submitted_at, repo_url, demo_url, tags, duplicate_of
assignmentsassignmentassignment_id, judge_email, project_id, project, status, source, created_at
scoresreview, drafts includedreview_id, judge_email, project_id, project, submitted_at, one column per criterion, weighted_0_1, comment
resultsproject in the latest runrun, rank, rank_low, rank_high, raw_rank, project_id, project, reviews, raw_mean, calibrated, se, excluded, input_digest
judgesjudge in the latest runrun, judge_email, reviews, offset, scale, noise, flag
auditaudit row for the eventseq, ts, actor, action, object_type, object_id, ip, before, after, prev_hash, row_hash

The file is served as text/csv with a download name like sample-hack-2026-scores.csv. Any cell that starts with =, +, - or @ (and isn’t a number) gets a leading ', so a spreadsheet won’t run a project title as a formula. results and judges always describe the latest run, published or not; the run column says which.

The OpenAPI document and the reference page

GET /api/schema serves the OpenAPI 3.0 document, generated by drf-spectacular from the same views, so it can’t drift from the code. It’s YAML by default; ask for JSON with ?format=json or Accept: application/json.

curl http://localhost:8080/api/schema                 # YAML
curl "http://localhost:8080/api/schema?format=json"   # JSON

GET /api/docs is a reference page rendered on the server from that same document. It needs no JavaScript and no CDN, so it works on a stack with no network. Both are public.

The schema lists each route’s success response only. The error codes are in this chapter and in each route’s description.

Running it for real

The image that runs the demo is the image you run an event on. What changes is configuration: a real secret, your hostname, no demo accounts, TLS in front, and backups. This chapter covers each, then the maintenance commands and the few sharp edges I know about.

A checklist

For an event that people outside your laptop will use:

  1. Start from an empty database, with the demo accounts off (below).
  2. Put your settings in a docker-compose.override.yml (below) rather than editing docker-compose.yml.
  3. Create the first site admin with createsuperuser.
  4. Put a reverse proxy with TLS and rate limits in front, and set DJANGO_SECURE=1.
  5. Schedule pg_dump and verify_audit.
  6. Before the event, run through the tour once on your own deployment.

Environment variables

Settings come from the environment, read in src/ballotbench/settings.py, src/entrypoint.sh and the seed command. This is all of them.

VariableDefaultIn docker-compose.ymlWhat it does
DJANGO_SECRET_KEYnonenot setsigns sessions and CSRF tokens; wins over the file below
DJANGO_SECRET_KEY_FILEnone/data/secret_keya file holding the key; the entrypoint writes a random one there on first boot if it’s missing or empty
DJANGO_DEBUGoffnot set1 shows Django’s debug pages and allows a built-in insecure key; for development only
DJANGO_ALLOWED_HOSTSlocalhost,127.0.0.1,[::1]adds webcomma-separated host names the portal answers to
DJANGO_SECUREoffnot set1 behind a TLS proxy: secure cookies, HTTPS redirect, HSTS, trust the proxy’s scheme header (below)
POSTGRES_DBballotbenchnot setdatabase name
POSTGRES_USERballotbenchnot setdatabase user
POSTGRES_PASSWORDballotbenchballotbenchdatabase password
POSTGRES_HOSTlocalhostdbdatabase host
POSTGRES_PORT5432not setdatabase port
BALLOTBENCH_DEMO_SEEDoff"1"1 loads the fixture event, the open demo event and the demo accounts with their fixed tokens on boot; anything else loads nothing
BALLOTBENCH_FIXTURES/app/fixtures.json in the imagenot setthe fixture file the seed imports on every boot
BALLOTBENCH_TRUSTED_PROXIES0not sethow many reverse proxies sit in front; the caller’s address is then read from X-Forwarded-For, counting from the right
BALLOTBENCH_ADDRESS_LIMIT_SCALE10not setmultiplies every per-address rate limit, so a venue behind one NAT address isn’t locked out; per-account limits aren’t scaled
DJANGO_EMAIL_BACKENDthe database outboxnot setDjango’s SMTP backend (django.core.mail.backends.smtp.EmailBackend, plus the usual EMAIL_* settings) to really send mail
WEB_WORKERS3not setgunicorn worker processes

If neither DJANGO_SECRET_KEY nor DJANGO_SECRET_KEY_FILE is set, and debug is off, the portal refuses to start rather than run with a guessable key. The generated key lives on the webdata volume, so sessions survive a restart and the key never lives in the repository.

The db service has its own POSTGRES_DB, POSTGRES_USER and POSTGRES_PASSWORD, read by the Postgres image. The web service’s values must match them. The Postgres image only reads them when it creates the database, the first time the pgdata volume is used; changing the password later means ALTER USER inside the database as well.

DJANGO_ALLOWED_HOSTS must keep 127.0.0.1: the container’s health check asks for http://127.0.0.1:8080/projects, and Django refuses a host it doesn’t know.

An override file

Compose reads docker-compose.override.yml next to docker-compose.yml automatically, so your settings stay out of the file you pull updates into. For a deployment at judging.example.org behind a proxy on the same machine:

# docker-compose.override.yml
services:
  db:
    environment:
      POSTGRES_PASSWORD: a-long-random-password
  web:
    environment:
      POSTGRES_PASSWORD: a-long-random-password
      DJANGO_ALLOWED_HOSTS: judging.example.org,localhost,127.0.0.1
      DJANGO_SECURE: "1"
      BALLOTBENCH_DEMO_SEED: "0"
    # Only the proxy on this machine talks to the portal.
    ports: !override
      - "127.0.0.1:8080:8080"
    # With DJANGO_SECURE=1 a plain-http request is redirected to https,
    # so the health check has to say it came through the proxy.
    healthcheck:
      test: ["CMD", "python", "-c", "import urllib.request as u; u.urlopen(u.Request('http://127.0.0.1:8080/projects', headers={'X-Forwarded-Proto': 'https'}), timeout=3)"]

!override replaces the port list instead of adding to it; it needs Docker Compose 2.24 or newer. Check what Compose will actually run with docker compose config.

Demo accounts

With BALLOTBENCH_DEMO_SEED=1, every boot makes sure these exist: the site admin admin@ballotbench.local, the organizer organizer@ballotbench.local, two fixture judges and a fixture participant, all with the password ballotbench-demo, and four API tokens whose values are printed in the README. Everything about them is public, so a real deployment must not have them.

  • On a new deployment, set BALLOTBENCH_DEMO_SEED to 0 before the first boot. None of them is created, including the admin, so create your own with createsuperuser.
  • If they already exist, setting the variable to 0 stops the seed creating them; it doesn’t delete them. Start again from an empty volume (docker compose down -v, which deletes everything), or switch them off from a shell:
docker compose exec web python manage.py shell -c "
from django.utils import timezone
from portal.models import ApiToken, User
ApiToken.objects.filter(label='demo', revoked_at=None).update(revoked_at=timezone.now())
User.objects.filter(email__in=['admin@ballotbench.local', 'organizer@ballotbench.local',
    'diego.herrera@example.org', 'ines.rocha@example.org', 'priya1@example.org']).update(is_active=False)
"

A deactivated user can’t sign in, and their tokens are refused.

The seeded events follow the same switch. With BALLOTBENCH_DEMO_SEED set to anything but 1, the seed loads nothing: no fixture event, no demo event, no accounts. A deployment that already has them keeps them (the seed never deletes); remove them for good with docker compose exec web python manage.py delete_event sample-hack-2026 --yes (and demo-open), and they won’t come back.

TLS and a reverse proxy

The portal speaks plain HTTP on port 8080. For anything public, put a reverse proxy in front that terminates TLS, and set DJANGO_SECURE=1. That turns on:

  • SESSION_COOKIE_SECURE and CSRF_COOKIE_SECURE: cookies only over HTTPS;
  • SECURE_PROXY_SSL_HEADER = ("HTTP_X_FORWARDED_PROTO", "https"): a request counts as HTTPS when the proxy says so in X-Forwarded-Proto;
  • SECURE_SSL_REDIRECT: plain-HTTP requests are redirected to HTTPS;
  • SECURE_HSTS_SECONDS of 30 days: browsers stay on HTTPS;
  • CSRF_TRUSTED_ORIGINS: https:// plus each host in DJANGO_ALLOWED_HOSTS.

It’s off by default so the demo works on http://localhost. With it on, the acceptance checker and the isolation probe, which talk plain HTTP to port 8080, get redirects instead of answers; they’re for the demo stack.

Trusting X-Forwarded-Proto is only safe if the proxy always sets it itself and nobody can reach port 8080 except through the proxy. Bind the port to 127.0.0.1 as in the override above, or keep it off the host entirely.

The portal limits logins per address and per account, and sign-ups and password resets per address (in the database, so every worker shares the counts; see BALLOTBENCH_ADDRESS_LIMIT_SCALE). A proxy is still the right place for a coarse limit on everything, before a request reaches Python. An nginx example:

limit_req_zone $binary_remote_addr zone=bb_auth:10m rate=10r/m;
limit_req_zone $binary_remote_addr zone=bb_api:10m rate=10r/s;

server {
    listen 80;
    server_name judging.example.org;
    return 301 https://$host$request_uri;
}

server {
    listen 443 ssl;
    server_name judging.example.org;
    ssl_certificate     /etc/letsencrypt/live/judging.example.org/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/judging.example.org/privkey.pem;
    client_max_body_size 1m;

    proxy_set_header Host $host;
    proxy_set_header X-Forwarded-Proto $scheme;
    proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;

    location ~ ^/(login|signup)$ {
        limit_req zone=bb_auth burst=5 nodelay;
        proxy_pass http://127.0.0.1:8080;
    }
    location /api/ {
        limit_req zone=bb_api burst=20;
        proxy_pass http://127.0.0.1:8080;
    }
    location / {
        proxy_pass http://127.0.0.1:8080;
    }
}

The limits are a starting point: ten sign-in or sign-up attempts a minute per address is plenty for a person and slow for a script. A whole venue behind one NAT address shares that budget, so watch the proxy’s log on the day.

The audit log records the address a request came from, as the portal sees it. It deliberately doesn’t trust X-Forwarded-For, so behind a proxy the address column shows the proxy’s address, not the visitor’s; the proxy’s own access log has the real one.

Email

The portal sends no email. Invitation links are shown once, to whoever makes them, to pass on however they like; that’s what lets it run with no network. EMAIL_BACKEND is Django’s console backend, so anything Django itself tries to send is written to the web container’s log.

If you add email, it’s a change in settings.py, since there are no email variables today. Replace the EMAIL_BACKEND line with something like:

EMAIL_BACKEND = "django.core.mail.backends.smtp.EmailBackend"
EMAIL_HOST = env("EMAIL_HOST", "localhost")
EMAIL_PORT = int(env("EMAIL_PORT", "587"))
EMAIL_HOST_USER = env("EMAIL_HOST_USER", "")
EMAIL_HOST_PASSWORD = env("EMAIL_HOST_PASSWORD", "")
EMAIL_USE_TLS = True
DEFAULT_FROM_EMAIL = env("DEFAULT_FROM_EMAIL", "ballotbench@judging.example.org")

and add those variables to your override file. Keep in mind that the stack then needs the network to reach the mail server.

Backups

All state is in two Docker volumes: pgdata (the database) and webdata (the generated secret key). The database is plain Postgres, so pg_dump works. Take a custom-format dump from the running stack:

docker compose exec -T db pg_dump -U ballotbench -Fc ballotbench > ballotbench-$(date +%Y%m%d-%H%M).dump

-T matters: without it Compose allocates a terminal and can mangle the binary output. I’d run this from cron every hour during an event, and keep the files off the machine.

To restore, stop the portal, recreate the database empty, load the dump, start the portal and check the audit chain:

docker compose stop web
docker compose exec -T db dropdb -U ballotbench ballotbench
docker compose exec -T db createdb -U ballotbench ballotbench
docker compose exec -T db pg_restore -U ballotbench -d ballotbench --no-owner < ballotbench-20260929-1800.dump
docker compose start web
docker compose exec web python manage.py verify_audit

Restore into an empty database, not over the live one. pg_restore loads the rows before it creates the triggers, so the audit chain, the deadline and the other rules come back exactly as they were, and verify_audit should report the chain intact. Losing webdata only signs everyone out; the entrypoint makes a new key.

Upgrades

docker compose exec -T db pg_dump -U ballotbench -Fc ballotbench > before-upgrade.dump
git pull
docker compose up -d --build
docker compose logs -f web

There’s no separate migration step. Every boot runs manage.py migrate --noinput, then the seed, then gunicorn, so a new version’s migrations (including new triggers) are applied when the new container starts; with no new migrations it’s a no-op. If an upgrade goes wrong, restore the dump you just took rather than trying to migrate backwards.

The Postgres image is pinned (postgres:18.4). A patch release is a change of tag; a new major version needs a dump and a restore into a fresh volume, as with any Postgres.

Management commands

Run them with docker compose exec web python manage.py <command>.

seed

Imports the fixture event and creates the open demo event, and with BALLOTBENCH_DEMO_SEED=1 the demo accounts. The entrypoint runs it on every boot. It’s idempotent: imported rows are keyed by their fixture ids, so a second run creates nothing and never overwrites what people changed. You shouldn’t need to run it by hand.

To load another event in the fixture’s shape, use the importer from a shell and give it its own slug:

docker compose cp other-event.json web:/tmp/other-event.json
docker compose exec web python manage.py shell -c "
import json
from portal.importer import import_event
event, counts = import_event(json.load(open('/tmp/other-event.json')), slug='other-event')
print(event.slug, counts)
"

seed --fixtures other-event.json looks like the way to do this, but it names every imported event sample-hack-2026, so it fails once the fixture event exists.

createsuperuser

Creates a site admin, who can do everything an organizer can in every event and use the Django admin at /admin/. You need one when the demo accounts are off.

docker compose exec web python manage.py createsuperuser

It asks for an email address and a password. From a script:

docker compose exec -e DJANGO_SUPERUSER_PASSWORD='a long passphrase' web \
  python manage.py createsuperuser --noinput --email you@example.org

verify_audit

Recomputes the audit log’s hash chain from the first row to the last and names the first row that doesn’t fit: a missing row, a changed prev_hash, or a row edited after it was written. It exits with status 1 if the chain is broken, so it can run from cron or CI.

audit chain intact: 26 rows, head e1cd7b2acec59727

Run it after a restore, before publishing results, and on a schedule. The database refuses edits to the log from the app and from ordinary SQL; this catches the one thing that can get past that, a database superuser.

normalization_proof

Recomputes every calibration claim in the method chapter from an event’s scores: which judges carry no weight and why, the agreement test, the ranking with its intervals, the invariance checks on these scores, and a benchmark on synthetic events. It takes a few seconds and is deterministic.

docker compose exec web python manage.py normalization_proof                       # the fixture event
docker compose exec web python manage.py normalization_proof --event your-event    # yours

Use it when you want to see the model at work on your own event before you publish, or to check the book’s numbers. It needs an event with reviews.

delete_event

Deletes an event and everything in it. After submissions close, the database refuses to delete submitted projects, which is what you want during an event and in the way afterwards; this command is the one deliberate way round it. It asks for --yes, and it writes an event.delete row to the audit log first. The audit rows of the deleted event stay, since they aren’t tied to it by a foreign key.

docker compose exec web python manage.py delete_event old-hack-2025 --yes

The maintenance flag

Some database rules would stop legitimate maintenance. The fixture’s projects were submitted before a close date that has already passed, so importing them is, to the deadline trigger, a late submission. Deleting a closed event means deleting projects after the deadline. The fixture’s team sizes aren’t checked against the event’s limit either.

For those cases the triggers look at a session setting, ballotbench.import. When it’s on, the deadline trigger and the team-size check step aside. Only two pieces of code set it, the fixture importer and delete_event, and both use SET LOCAL, so it lasts only until their own transaction ends, and both write an audit row saying what they did.

It’s not a security boundary. Anyone with SQL access to the database can set it, or drop a trigger outright. The triggers are there to stop the app, the admin and a careless shell from breaking the rules by accident. Deliberate changes by someone with database access are what the audit chain and verify_audit are for. The app connects as the user the Postgres image creates, which is a superuser, so keep database access to the people who run the event.

API tokens

Anyone signed in issues and revokes their own tokens at /me/tokens; a token acts with that person’s roles and nothing more, is shown once, and only its hash is stored. An admin can also issue one for any account from a shell:

docker compose exec web python manage.py shell -c "
from portal.auth import issue_token
from portal.models import User
print(issue_token(User.objects.get(email='organizer@example.org'), 'results script'))
"
bb_YC5-kg1ErFIezXIBwN0h1Y6ysq9uxVNkEDkaB2_IEmM

The label is for you, to tell tokens apart. The token has all of its user’s roles. To revoke it, set its revoked time, from the Django admin (Api tokens) or from a shell:

docker compose exec web python manage.py shell -c "
from django.utils import timezone
from portal.models import ApiToken
ApiToken.objects.filter(user__email='organizer@example.org', label='results script',
                        revoked_at=None).update(revoked_at=timezone.now())
"

A token issued or revoked from a shell isn’t in the audit log; one revoked through the admin is.

With the network off

Once the images are built, the stack needs no network. The override docker-compose.offline.yml puts both containers on a Docker network with no route out:

docker compose down -v
docker compose -f docker-compose.yml -f docker-compose.offline.yml up

Docker doesn’t publish ports from an internal network, so in this mode the portal isn’t reachable from the host’s browser. Check it from inside, as CI does:

docker compose exec web python -c "import urllib.request as u; [u.urlopen('http://localhost:8080' + p) for p in ('/projects', '/static/portal/site.css')]"

docker compose up --force-recreate puts the containers back on the normal network.

Known sharp edges

Things I’d want to know before running an event on it:

  • An account must confirm its address before it can vote, but anyone with a working inbox can sign up. Logins, sign-ups and resets are rate limited; a proxy with its own limits is still wise.
  • Behind a proxy, set BALLOTBENCH_TRUSTED_PROXIES to the number of proxies, or the audit log and the rate limits see the proxy’s address.
  • Mail goes to the outbox table, readable in the admin. For real delivery, set DJANGO_EMAIL_BACKEND to Django’s SMTP backend and configure it.
  • Calibration runs inside the organizer’s request. On the fixture that’s a few seconds; calibration needs no background worker.
  • Community vote tallies are computed when read. Voiding a ballot after publishing changes the public numbers (and is in the audit log).

The acceptance checker

Dogfood’s organizers give every team the same checker, run.py: a standard-library Python script that reads a team’s .dogfood.toml, makes seven HTTP requests against the running portal, and prints a report of which tiers it could verify. This chapter explains the contract as ballotbench meets it, and the stricter probe I wrote to go past it.

.dogfood.toml, line by line

[portal]
base_url = "http://localhost:8080"

[tiers]
claimed = ["T1", "T2"]
pitch = "A self-hosted hackathon portal whose judging you can defend: weighted rubrics, calibrated judges, and deadlines and isolation enforced in the database."

[auth]
organizer   = "Authorization: Bearer bb_demo_organizer_5c1e0a"
judge_a     = "Authorization: Bearer bb_demo_judge_a_8d24f1"
judge_b     = "Authorization: Bearer bb_demo_judge_b_3a9e77"
participant = "Authorization: Bearer bb_demo_participant_61b0c4"

[routes]
gallery      = "/projects"
submit       = "/api/events/sample-hack-2026/projects"
judge_scores = "/api/judge/scores"
peer_scores  = "/api/judge/scores?judge=jdg_24"
csv_export   = "/api/events/sample-hack-2026/export/scores.csv"
LineMeaning
base_urlwhere the checker sends every request: the port docker compose up publishes
claimedthe tiers this entry claims: T1 (submissions and the gallery) and T2 (judging and isolation), the ones run.py has checks for. The public vote (T3) and the stretch pieces (T4) are built and tested, but the checker can’t verify them, so they aren’t claimed.
pitcha one-line description of the entry; run.py doesn’t read it
organizer, judge_a, judge_b, participanta complete HTTP header for each role. The checker splits it at the first : and sends it as is. These are the demo tokens the seed creates when BALLOTBENCH_DEMO_SEED=1.
gallerythe public page listing submitted projects
submitwhere a participant creates a project, here in the fixture event, which is closed
judge_scoreswhere a judge reads their own scores
peer_scoresthe URL that would return judge A’s scores: jdg_24 is judge A’s id in the fixture
csv_exportan organizer’s CSV export

The four accounts are chosen so the checks mean something: judge_a (jdg_24) and judge_b (jdg_29) both have fixture reviews and share no project, so each has scores the other must not see, and the participant is on a real fixture team (NorthKiln), so the late submission is refused for being late, not for having no team.

The seven checks

TierCheckThe requestPasses whenWhat answers it
T1gallery is publicGET /projects, no header200the gallery view, which lists submitted projects to anyone
T1project from fixtures shownthe same responseit contains the title of one of the fixture’s first three projectsthe seed imports the fixture on boot, and the gallery orders page one by id, so Glass Signal, Small Meadow and Deep Compass lead
T1closed event refuses submissionsPOST to submit as participant, with a title and summaryany 4xx409 with the close time: the API checks the window on the database’s clock, and the project_deadline trigger would refuse it anyway
T2judge sees own scoresGET judge_scores as judge_a200/api/judge/scores returns the caller’s own reviews
T2judge cannot see peer scoresGET peer_scores as judge_b401 or 403403: a judge may only name themselves, and gets a refusal, never an empty list
T2participant blockedGET judge_scores as participant401 or 403403 judges only
T2csv export worksGET csv_export as organizer200, and the first line has a commathe scores export, whose first line is the CSV header

The checker is lenient in places (any 4xx for the late submission, 401 or 403 for the peer probe). ballotbench answers with the specific code in each case; the API chapter has the contract.

A tier counts as verified only if every one of its checks passes and every tier below it is verified too. run.py has checks for T1 and T2 only, so those are the most it can verify.

Regenerating acceptance-report.txt

The report in the repository is the checker’s output against the Docker build on a fresh volume. To make it again, from the repository root:

docker compose down -v
docker compose up -d --build --wait
python3 run.py .dogfood.toml > acceptance-report.txt

Run it from the root so it finds fixtures.json (it looks in the current directory, next to run.py, next to the config, and in a data/ folder beside the config; --fixtures path names it outright). Any Python 3 works: 3.11 and newer read the TOML with tomllib, older ones with the script’s own small parser. The last line is the verdict:

claimed T1 T2, verified T1 T2

CI does the same on every push to main and fails if that line is anything else. The checker writes nothing to the portal apart from the audit row every CSV export leaves; its one write request is refused.

The isolation probe

The checker tries one peer probe and one participant probe. Passing it means little on its own: a portal that returned 403 for every judge request would pass. So scripts/isolation_curl.sh tries 94 things over HTTP, each with the exact status it must get back:

sh scripts/isolation_curl.sh                         # against http://localhost:8080
sh scripts/isolation_curl.sh http://localhost:9000   # or another base URL

It needs sh, curl and sed. First it reads the ids it needs with their owners’ own tokens (one of judge A’s assignments, a project judge A reviewed, the participant’s project, another team’s project), then:

GroupAttemptsExpected
judge scoresjudge B names judge A by fixture id and by email, and names the constant judge; the participant reads scores and lists assignments; no token; a made-up token; judge A reads their own403, 401, and 200 for the last
someone else’s reviewjudge B and the participant post scores to judge A’s assignment; judge B opens judge A’s scoresheet page404
deadlinethe participant submits and edits after the close; edits another team’s project; a judge and an anonymous caller create projects409, 403, 401
exportsall eight kinds, anonymously, as the participant, as a judge, and as the organizer401, 403, 403, 200
results before publicationanonymous, judge and participant read the fixture’s results404
organizer pagesa judge’s and the participant’s tokens on every organizer page403
community vote and commentsreading and casting ballots with no or a made-up token, voting where there’s no vote, a made-up voting link, unpublished results, commenting anonymously401, 404
signed records, bundles and webhooksjudge B fetching judge A’s record or certificate, the participant and anonymous callers fetching records, exporting or importing bundles, opening the webhooks page; judge A fetching their own record403, 404, 401, and 200 for the last
public pagesthe gallery, a submitted project and its comments, the API schema, the signing key, the embeddable gallery200

It prints one line per attempt and a total, and exits with status 1 if anything came back different:

== judge scores
ok   200 judge_a reads own scores
ok   403 judge_b names judge_a by fixture id
...

94 of 94 as expected

Two things to know when running it:

  • It expects the fixture’s results to be unpublished. If you’ve published them on the stack, the three results lines fail, correctly. Reset with docker compose down -v.
  • It speaks plain HTTP to port 8080, so it’s for the demo stack. A deployment with DJANGO_SECURE=1 redirects it to HTTPS.

Like the checker, a passing run changes nothing: every write it tries is refused. The organizer’s exports leave their audit rows.

The test suite goes further than both (the same rules through the pages, the API, the admin and raw SQL against a real Postgres), but the probe is the one to run against a deployment, because it asks the running thing.

Build notes

A running log of what surprised me, what I changed my mind about, and why. Newest last.

Team names repeat in the fixture

I started with UNIQUE (event, name) on teams, which felt obviously right. The seed fell over on its first run: the fixture has three different teams called StillTrail and two each called AmberSwitch and OpenSignal. They have different ids and different members, so they are different teams. Dropped the constraint; teams are told apart by id, and the UI shows the id-based link. Projects, by contrast, are checked for duplicates on purpose.

The planner’s first version wasn’t balanced

My first planner went project by project, giving each the least-loaded eligible judge. A property test with 12 projects, 6 judges and k=3 ended at loads 7, 7, 6, 6, 5, 5, not 6 each. Driving by judge instead (the idlest judge picks the neediest project) didn’t fix it either: by the end, the two idle judges already held both of the last projects that needed someone. Greedy can’t see that coming. What fixed it was a repair pass afterwards: move one of the plan’s new reviews from the busiest judge to the idlest one who is allowed to take it, until no move narrows the gap. Existing reviews are never moved. The property test now checks loads stay within one on 200 random events.

My first voting link confirmed the voter on GET, which is what my plan said and what every tutorial does. Then I remembered that corporate mail scanners (and some webmail previews) fetch every link in a message before the person sees it. With a single-use token, the scanner would use it up and the voter would click a dead link. Opening the link now shows a button, and the POST behind it confirms. One extra click, and the link survives being looked at.

A refused request that forgot it happened

The rate limiter counts hits in Postgres, inside the request’s transaction (every request is atomic). My API views raised Conflict for an over-budget ballot, DRF’s exception handler marks the transaction for rollback, and the rollback took the rate-limit hit with it. So failed requests were free, which is exactly backwards: the requests worth limiting are the ones that fail. The API views now return their refusals as responses instead of raising them, and the hit commits. The page views already did.

Where the ballot’s privacy stops

I went back and forth on what vote.cast should record. The choices in the clear would make the log a full record, but every organizer of the event reads that log, and a community ballot is between the voter and the tally. So the row has the voter, the credits spent, the number of projects and a fingerprint of the ballot. A plain SHA-256 of something as small as a quadratic ballot can be reversed by trying every ballot, so the fingerprint is an HMAC keyed with the server’s secret. It still lets the abuse panel see several voters casting the same ballot, which is the one thing it was for.

One inbox, one ballot, and not telling anyone who voted

My plan said a second sign-up with another spelling of the same inbox (a.b+x@googlemail.com after ab@gmail.com) must be refused. Saying “that inbox has already voted” would let anyone check whether a given address voted. Instead the page reads the same either way, no second ballot is made, the attempt is logged as vote.duplicate_refused for the abuse panel, and a fresh link goes to the address just typed. (At first it went to the address that signed up first; see the end of these notes for why not.) The person who owns the inbox can still get in; nobody else learns anything.

Hiding results while the vote is open

Refusing to publish while voting is open wasn’t quite enough: an organizer could publish the judges’ results first and then open a vote, and voters would have the ranking in front of them. Published results now go back to a 404 for the public while a vote is open, and come back when it closes.

Tier T4: what surprised me

A route would have swallowed another. I first added api/events/import at the end of urls.py, but the older api/events/<slug:slug> matches first, with slug import, and would have answered every import with 405. It sits above that pattern now, with a comment saying why.

Pasted JSON can’t always be echoed back. The verify API returns the record it was sent, so a record containing a lone surrogate ("\ud800", perfectly legal JSON) made DRF’s renderer throw a 500. The malformed-input test found it. Looking for its siblings turned up NaN, which Python’s json reads but DRF’s strict output refuses. Both now fail the canonical encoding step and come back as “not valid”, and both are test cases.

SET LOCAL outlives the function that sets it. The importer turns on the deadline bypass for “its own transaction”, but inside a request with ATOMIC_REQUESTS its transaction is only a savepoint: the bypass would have stayed on for the rest of the request. Harmless for the seed, not for an API endpoint, so the import now switches it off before returning (a failed import’s savepoint rollback undoes it anyway).

The fixture importer trusted its file; a bundle importer can’t. Bundles come from any organizer, and the fixture path wrote repo_url straight into a link. A javascript: URL would have been stored XSS on the project page. Bundles now go through the same URL, email and choice checks the forms use.

The round trip calibrates identically, not approximately. I expected to need a tolerance: floating point sums depend on order. But the export lists reviews by id and the import creates them in that order, so the fit sees the same sequence and every q matches to the last bit. The test asserts equality.

An iframe that reports scrollHeight can grow but never shrink: document.documentElement.scrollHeight is never less than the iframe’s own height. The embed reports the height of its content box instead; I checked it in headless Chromium on a page that embeds the widget.

Three branches at once, and what the docs caught

I split the last stretch into three branches built side by side: the book, the public vote, and the T4 pieces. Two things I didn’t expect:

Both code branches took migration 0013. Each was right on its own branch and together they gave Django two leaf nodes. The webhooks migration became 0014 and depends on the voting one; nothing outside a scratch database had applied either, so renaming was safe. Next time I’d reserve numbers up front.

Writing the docs found real bugs. Explaining the calibration page line by line turned up that the page and normalization_proof disagreed: the page leaves out a project whose every reviewer the model ignores (prj_24, both of whose judges are discordant), the proof didn’t, so JUDGING.md quoted a ranking the product never shows. They also ran the agreement test with different shuffle counts. Documenting the operations side found that BALLOTBENCH_DEMO_SEED=0 still loaded the fixture event, that DJANGO_SECURE made the health check fail on its own redirect, and that an organizer’s ?judge=jdg_24 could match a judge in another event, since fixture ids are only unique per event. All fixed, each with a test. The lesson I keep relearning: the fastest code review is trying to explain the code to a stranger.

What a security review found in the vote and the duplicates

Void, read, unvoid. Organizers saw the tallies live and could undo a void, so voiding one ballot, reading the tallies and counting it again showed exactly what that voter chose, while every page said ballots were private. A void is now final, and while voting is open the organizer page shows how many ballots there are, not the tallies. After the close the tallies are there, and a void still subtracts a ballot you can see; but it is gone for good, so reading it costs the organizer that vote.

The first spelling of an inbox owned its links. Asking for a link as alice+nope@yahoo.com sent every later link for alice@yahoo.com to the +nope address. For providers where a +tag is its own inbox, that’s a stolen ballot. The link now goes to whatever was typed; a new link replaces the old one. Quoted local parts ("a.b"@gmail.com) are refused: legal, but nobody needs one to vote, and normalizing them right is a rabbit hole. normalize_email also used to raise on +x@example.com, and because it runs over every team member’s address, one such account broke every email voter’s ballot page. It no longer raises.

Duplicates by pk. The detector ordered projects by pk, drafts included, and every submit re-flagged the whole event. A draft made early and filled in later with another team’s public title and repo made the real project the “duplicate”, taking it out of judging and off the ballot; and an organizer’s “it’s different” lasted until the next submit. Now only submitted projects count, earlier means submitted_at, a submit only flags the project being submitted, and duplicate_cleared keeps an organizer’s call.

What a stranger found in an hour

With every test green, I had someone follow the book’s tour from a fresh clone, as a judge would, and write down everything that disagreed with it. It found 28 things. The ones that stung:

  • “Save draft” after submitting quietly replaced the submitted scores. The review stayed “submitted”, with the old timestamp, and calibration read the new numbers. Every test submitted once and stopped. A submitted review now only resubmits.
  • Deleting a rubric criterion after scoring was a 500. Changing a weight was refused politely by the trigger; deleting went through a foreign key set to RESTRICT, and nobody had wrapped that path. The same shape as the track bug the security review found an hour earlier: a database refusal reached from a path the app didn’t guard.
  • The venue problem. Rate limits per address are right against a flood and wrong at a hackathon, where three hundred people share one NAT address. The walkthrough locked itself out after twenty ordinary logins. Per-address limits now scale for a crowd, and a per-account limit does the job of stopping someone guessing one person’s password.
  • The offline proof published no port. internal: true networks can’t publish ports, so following the README gave a stranger nothing to open. The honest instruction is simpler: switch off the Wi-Fi.

None of these were security holes and all of them would have been on camera in a demo. Reading the docs aloud against the running thing is a test nothing else replaces.