ballotbench
ballotbench is a hackathon portal you run yourself. Teams submit projects,
judges score them against a rubric the organizer weights, and the organizer
gets a ranking that takes each judge’s habits out and says honestly how sure
it is. It starts with one docker compose up and needs no network after
that.
I built it for Dogfood 2026, a hackathon where the thing you build is a hackathon platform, and the winner’s portal gets used to run the next events. That shaped every decision: it has to come up on a volunteer’s laptop, and its results have to survive a team asking “why did we come fourth?”
What makes a ranking defensible
Three things, which the rest of this book keeps coming back to:
- Nobody can see what they shouldn’t, or change what they shouldn’t. Judges can’t read each other’s scores; nobody can edit a project after the deadline. These rules live in the database, not only in the pages, so they hold for every way in: the web pages, the API, the admin, a shell.
- The maths is written down and tested. Judges score differently; the calibration model that corrects for it is described in full, and its guarantees (shifting or stretching one judge’s scores changes nothing, a judge who gives everyone the same score counts as absent) are checked by tests and by a command you can run.
- Everything leaves a trail. Every change writes a row to an audit log that the database won’t let anyone edit, chained by hashes so a quiet edit by a database superuser still shows.
Where to start
-
To see it first: the five-minute demo video.
-
To try it: Run it in five minutes.
-
To see a whole event: A whole event, start to finish.
-
To judge the judging: Assignment, scoring and calibration.
-
To run it for a real event: Running it for real.
The source is at https://github.com/keirsalterego/ballotbench, MIT licensed.
This book
This book is published at https://keirsalterego.github.io/ballotbench/,
rebuilt whenever the docs change on main. Its source is in docs/, and the
chapters on judging, architecture and the data model include the
repository’s top-level documents directly, so there is one copy of each. To
read it locally with mdBook 0.5:
mdbook serve docs --open
Run it in five minutes
You need Docker with the Compose plugin. Nothing else: no Python, no Postgres, no accounts anywhere.
git clone https://github.com/keirsalterego/ballotbench.git
cd ballotbench
docker compose up
The first run builds the image, starts Postgres, creates the tables, loads the shared Dogfood fixture event and prints something like:
fixture event sample-hack-2026: 8 tracks, 30 judges, 40 teams, 91 members, 41 projects, 126 reviews, 1 duplicate
seeded. test logins:
organizer Authorization: Bearer bb_demo_organizer_5c1e0a
judge_a Authorization: Bearer bb_demo_judge_a_8d24f1
judge_b Authorization: Bearer bb_demo_judge_b_3a9e77
participant Authorization: Bearer bb_demo_participant_61b0c4
Open http://localhost:8080. You’re looking at the public gallery.
Sign in
Every demo account has the password ballotbench-demo:
| Who | |
|---|---|
| Organizer | organizer@ballotbench.local |
| Judge A | diego.herrera@example.org |
| Judge B | ines.rocha@example.org |
| Participant | priya1@example.org |
| Site admin | admin@ballotbench.local |
The bearer tokens are for scripts and the API: pass them as an
Authorization header.
Two events are waiting
- Sample Hack 2026 is the fixture: 41 projects, 126 reviews, on its real dates, so submissions closed in March 2026. Sign in as the organizer, open it, and go straight to Calibration and results.
- Demo Hack (open) is empty and open for submissions for two weeks from your first boot. Use it to walk through an event yourself; the tour does exactly that.
Check what it claims
python3 run.py .dogfood.toml # the official Dogfood checker (T1, T2)
python3 scripts/check_t3_t4.py .dogfood.toml # the same style of checks for T3 and T4
sh scripts/isolation_curl.sh # tries to reach what it shouldn't
docker compose exec web python manage.py normalization_proof # the judging maths, recomputed
docker compose exec web python manage.py verify_audit # the audit log's hash chain
With the network off
The simplest proof: once the images are built, switch off your Wi-Fi (or
pull the cable) and run docker compose up. The portal comes up, seeded,
and http://localhost:8080 works, because a port on your own machine needs
no outside network.
For a stricter proof, with no route out of the containers at all, there’s an override. A network with no way out can’t publish a port to your machine either, so you check it from inside the container, which is what CI does on every push:
docker compose down -v
docker compose -f docker-compose.yml -f docker-compose.offline.yml up -d --wait
docker compose exec web python -c "import urllib.request as u; print(u.urlopen('http://localhost:8080/projects').status)"
docker compose down && docker compose up -d --force-recreate # back to normal
Start over
docker compose down -v deletes the database volume. The next up seeds a
fresh copy.
A whole event, start to finish
This walks the open demo event through its life: set up, teams, submissions, the deadline, judging, calibration, results. It’s also the script of the five-minute demo video. Keep two browser windows open (a normal one and a private one) so you can be two people at once.
1. The organizer sets it up
Sign in as organizer@ballotbench.local, open My events, then Manage
next to Demo Hack (open).
- Dates and settings. Four windows: submissions, judging, voting and
(implicitly) results.
k, reviews per project, defaults to 3. - Tracks and prizes. Add a track called Hardware; add a prize for it.
- Rubric. Three criteria to start: functionality, quality, innovation, each 1 to 5, weight 1. Give functionality weight 2. The weights can change freely until the first score comes in; after that the database refuses, so nobody can tune weights after seeing who they favour.
- Invite a judge. Under Invite judges and organizers, pick Judge, tick a track, and make a link. It’s shown once and works once. The portal sends no email; send the link however you like.
2. Participants form a team
In the private window, create an account, then My events → join as a participant on the demo event.
- Create team, then Make an invite link. Open it signed in as someone else and they join. The link then stops working; a team can’t grow past the event’s size limit, even if two people accept at the same moment.
- Start your project. Save it as a draft: only the team and the organizers can see it. Submit it and it appears in the public gallery. You can keep editing until the deadline.
3. The deadline
As the organizer, move Submissions close to a minute from now and save. Wait for it. Back in the participant’s window, try to save the project: the answer is That window is closed, with the time. The same happens through the API (a 409 with the close time) and in the admin. Behind all three, a trigger in the database refuses the write on its own clock, so even a bug in the app couldn’t let a late edit through.
4. Judging
As the organizer, open Hand out reviews. The preview shows which judge gets which project: every project gets k reviews, the idlest judge picks next, nobody reviews their own team or outside their tracks. Apply it.
Sign in as the judge. Your reviews lists only your projects and how long they’ll take. Each has a scoresheet: one bubble per score per criterion, a comment only the organizers see, Save draft and Submit review (after submitting, only Resubmit). If you know the team, I have a conflict steps you aside; the organizer sees the project needs another judge.
Meanwhile, the organizer’s Progress page shows each judge’s assigned, started and finished counts and which projects are short of reviews. Keep this page up to date refreshes it every 20 seconds.
5. Calibration
Open Calibration and results and run it. For the fixture event this is where it gets interesting:
jdg_07gave every project 4 on everything. They’re listed as gave every project the same score and carry no weight.jdg_01andjdg_23wrote one review each: nothing to compare with, so no weight either.- The ranking shows raw rank, calibrated rank, and the range each rank could plausibly be. On the fixture these ranges are wide, and the page says why: the judges agree no more than chance would.
Reading the calibration page goes through it line by line.
6. Publish
Publish results freezes the ranking to this calibration run. Until then, the results page and API answer 404 to everyone but the organizers. The public page shows each project’s calibrated score, its plausible rank range, and the fingerprint of the scores it came from.
7. A community vote
In Settings, set Voting mode to signed-in accounts and put the voting window after submissions close (the portal refuses it earlier). While it’s open, Community vote shows how many ballots are in and who cast them, with flags for crowded networks, brand-new accounts and identical ballots, but not the per-project tallies: those appear once voting closes.
A voter opens Vote on the event: the projects come in an order made for them, they spread 25 credits (n votes on a project cost n²), and they can’t vote for their own team. An account has to confirm its email address first: the link is sent at sign-up, and with the network off it lands in the outbox, which the site admin reads under Admin → Outbound emails.
Results can’t be published while the vote is open. Close it, publish, and the public results show the community votes next to the judges’ ranking.
8. Export
Exports has a CSV for every stage: registrations, teams, projects, assignments, every score, the calibrated results, each judge’s calibration, and the audit log with its hash chain. Every download is itself in the audit log.
The five-minute demo
The automatic version
scripts/demo.sh runs a whole event against the portal and narrates it,
printing the page to show in the browser at each step:
From a fresh clone, nothing else to set up:
git clone https://github.com/keirsalterego/ballotbench.git && cd ballotbench
sh scripts/demo.sh # builds and starts the portal if it isn't running, then runs
sh scripts/demo.sh --fresh --pause # empty database first, then wait for Enter before each step
BALLOTBENCH_PORT=9000 sh scripts/demo.sh # if port 8080 is taken
It creates a new event each run (Demo Day plus the time), so it can run again without a reset. In twelve steps:
- An organizer creates the event, with a track, a prize and a weighted rubric.
- Four judges accept single-use invites.
- Four teams sign up and submit; one invites a teammate.
- The deadline passes: a late edit is refused by the API and the page.
- The organizer hands out reviews.
- Four judges score: harsh, generous, steady, and one who gives everything a 3.
- Isolation attacks between judges and teams are refused.
- Calibration flags the constant judge and ranks with honest ranges.
- A community vote with a quadratic budget, each voter in their own order, and publishing refused while it’s open.
- Publishing, the signed results verified, a doctored copy refused.
- A team reads why it placed where it did; another team can’t.
- The exports and the audit trail.
It needs only Python 3 (standard library) and, for --fresh, Docker.
By hand
A shot list for recording one full event lifecycle (create, submit, judge, publish) in five minutes. It uses both seeded events: the open Demo Hack for the live lifecycle, and Sample Hack 2026 (the fixture, 126 real reviews) for judging maths worth showing.
Before recording:
docker compose down -v && docker compose up -d --build --wait
Open three browser windows (a normal one and two private ones) so you can be
three people at once. Every demo password is ballotbench-demo.
| Time | Who | What to do | What to say |
|---|---|---|---|
| 0:00 | anyone | Open http://localhost:8080. | One docker compose up, no network needed, seeded with the shared Dogfood fixture. |
| 0:15 | organizer@ballotbench.local | My events → New event, or open Demo Hack (open) → Manage. Show dates, add a track, set functionality weight 2. | Organizers set windows, tracks, prizes and a weighted rubric. The rubric freezes at the first score. |
| 0:45 | new account (private window) | Create an account, join as a participant on Demo Hack, Create team, Make an invite link. | Single-use invite links, stored only as hashes. |
| 1:05 | same | Start your project, fill title and repo, Submit. Show it in the gallery, then open Edit again and leave that tab open. | Drafts are private; submitted projects are public. |
| 1:25 | organizer | Settings: move Submissions close to a minute ago, save. | Watch the deadline hold on every path. |
| 1:35 | participant | In the edit tab you left open, press Save changes: That window is closed. Reload the event page: the project is now read-only. Then in a terminal, the fixture event, which closed in March: curl -X POST -H "Authorization: Bearer bb_demo_participant_61b0c4" -H "Content-Type: application/json" -d '{"title":"late"}' localhost:8080/api/events/sample-hack-2026/projects → 409. | The page, the API and a database trigger on the database’s own clock all refuse; even a bug in the app couldn’t let a late edit through. |
| 1:55 | terminal | python3 run.py .dogfood.toml | The official checker: T1 and T2 verified. |
| 2:05 | terminal | sh scripts/isolation_curl.sh (let it scroll) | 94 attempts, each answered with exactly the status it must get: 78 refused, 16 allowed (public pages and your own data). |
| 2:15 | diego.herrera@example.org (judge) | Your reviews (in the header) → open a project: the scoresheet with ballot bubbles. | Judges see only their own assignments; another judge’s scoresheet is a 404, naming another judge in the API is a 403. |
| 2:35 | organizer | Sample Hack 2026 → Manage → Hand out reviews: the preview proposes 8 top-ups. | The fixture’s two unfinished batches: the planner finds them and balances load. |
| 2:50 | organizer | Calibration and results → Run calibration. Scroll: flagged judges, does one judge decide the podium?, the ranking with could be ranges and the pairwise column. | The fixture’s constant judge (jdg_07, listed by name as Iva Petrova) gave every project 4: weight zero, exactly as if absent. Every rank comes with an honest range; on this data the judges agree no more than chance, and the page says so. |
| 3:35 | organizer | Publish results. Open the public results page; click signed copy; paste it into /verify: Valid. | Published results are frozen to one run, fingerprinted, and signed with Ed25519. |
| 3:55 | priya1@example.org (team NorthKiln) | Results → How your project was scored (Glass Signal, third). | “Why did we come third?” answered: each review, the judge’s habit, how much it counted. Judges stay anonymous. |
| 4:15 | organizer | Audit log: filter by results. Then Exports: download scores.csv. | Every change is in a hash-chained log the database won’t let anyone edit. Every stage exports as CSV. |
| 4:35 | terminal | docker compose exec web python manage.py verify_audit and ... normalization_proof | tail -8 | The chain checks out; the maths claims are recomputed live and hold. |
| 4:50 | Open the book: https://keirsalterego.github.io/ballotbench/ (it redirects to the custom domain). | MIT licensed; everything above is documented for a stranger. |
If there’s time to spare, show the community vote on Demo Hack: in Settings set Voting mode to signed-in accounts with a window that’s open now, then vote as priya (confirmed demo account) and show the quadratic budget and the organizer’s abuse panel.
For organizers
You organize an event if you created it, or someone sent you an organizer invitation for it. Site admins can do everything organizers can, in every event.
Creating an event
My events → New event. Anyone who already organizes an event, and any site admin, can create one; you become its first organizer. Every event starts with a three-criterion rubric you can change.
The slug is part of every URL and can’t change later. All times are UTC.
The windows
| Window | What it controls |
|---|---|
| Submissions open → close | Teams can form, and projects can be created, edited and submitted. Enforced by the database clock. |
| Judging open → close | Judges can save and submit reviews. |
| Voting open → close | Public voting, if you turn it on. Results can’t be published while it’s open. |
Moving a window takes effect immediately: extending the deadline reopens editing at once.
Tracks, prizes and the rubric
Tracks group projects and limit which judges see what: a judge with tracks only reviews projects in those tracks. A judge with no tracks reviews anything.
The rubric is a list of criteria, each with a weight and a range. A review’s score is the weighted mean of its criteria, each scaled to 0..1 by its own range. Change it freely until the first score arrives; after that the database freezes it.
Judges and co-organizers
Invite them with a link from the settings page. Links are single use and last seven days. A judge invitation can carry tracks. Nobody can be both a judge and on a team in the same event: the database refuses either order.
Handing out reviews
Hand out reviews shows a plan before it changes anything. It tops every submitted project up to k reviews, gives each new review to the idlest judge who may take it, and adds a few bridging reviews if some judges share no projects with the rest (calibration needs them connected). Projects no eligible judge is left for are listed, so you know to invite someone.
You can also give one project to one judge by hand on the Progress page, and take back a review nobody has started.
Watching progress
Progress lists each judge’s assigned, started, finished and stepped-aside counts, and each project’s finished reviews, fewest first. Projects below k are flagged low coverage.
The community vote
Turn it on in Settings with Voting mode: signed-in accounts (anyone
with an account votes once) or confirmed email addresses (anyone votes
once per inbox, after opening a link we mail them). Vote credits is each
voter’s budget: n votes for one project cost n² credits. Send people to
/events/<slug>/vote; the gallery and every project page link there too.
With the network off, confirmation links can’t leave the box. They land in Outbound emails in the admin, where a site admin can read them.
Community vote in the organizer menu shows how many ballots are in while voting is open, and the tallies once it closes (only organizers see them until you publish results, which you can’t do while voting is open), and a list of things worth a second look: several voters on one network, accounts made just before their ballot, identical ballots, and sign-ups refused because another spelling of the same inbox had already voted. None of these void anything by themselves. If you decide a ballot is fake, void it with a reason; it leaves the tallies for good and the audit log records it. There’s no undo: voiding a ballot and counting it again would show you what that voter chose.
Duplicates
When a project has the same repository and title as an earlier one (or the same team resubmits the same repository or title), it’s flagged and left out of the rankings. Only submitted projects count, and the one submitted later is the copy. Duplicates lets you confirm it or put it back; a project you put back is never flagged again.
Calibration and results
Run calibration as often as you like; each run is kept. Publishing freezes the public results to one run. You can hide them again. See Reading the calibration page.
The audit log
Audit log lists every change in the event: who, when, from which address,
and the before and after values. Filter by kind of change or by person. Rows
can’t be edited or deleted, even by the admin; verify_audit checks the hash
chain.
Exports
Every stage as CSV. Cells that a spreadsheet would run as a formula (a title
starting with =, say) are prefixed with a quote, so opening an export is
safe.
For participants
Joining
Create an account (Create an account, top right), then on My events choose join as a participant next to the event. If a teammate sent you an invite link, just open it and sign in; it adds you to their team.
Your team
You can be on one team per event. Start one from the event page, then Make an invite link for each teammate. Each link works once, and stops working after 72 hours or when submissions close, whichever comes first. You can revoke a link you haven’t used. The event sets the largest team size.
Your project
A team usually submits one project, but may submit more (the fixture has a team with two); each is judged on its own.
Start your project and fill in what you have: title, a one-line tagline, a summary, the description, links to the code and a demo, tags and a track.
- Save draft keeps it private to your team and the organizers.
- Submit puts it in the public gallery.
- After that, Save changes updates it; it stays submitted and public. You can keep editing until the deadline.
- A draft you don’t want can be deleted from its edit page (Delete this draft). A submitted project can’t be deleted by the team; ask the organizers.
Starting a second project with the same title as one your team already has takes you to the existing one instead, so a double click doesn’t leave two copies.
The deadline
At the close time, editing stops: the edit page shows your project read-only, and the page, the API and the database all refuse changes, on the server’s clock, not yours. If the organizers extend the deadline, editing reopens straight away.
Duplicates
If your project looks like one submitted earlier (same repository and title), it’s flagged for the organizers when you submit it, and left out of the rankings until they decide. Only the later submission is ever flagged, and once an organizer says it’s a different project, it stays that way. This catches accidental double submissions; if yours is flagged by mistake, tell the organizers.
For judges
An organizer sends you an invitation link. Open it, sign in or create an account, and accept. Your reviews then lists the projects assigned to you, and how long the rest will take at about ten minutes each.
Scoring
Each project has a scoresheet: the project’s description and links, then one row of bubbles per criterion. Pick a number in each row. The weight of each criterion is shown next to its name.
- Save draft keeps your scores without submitting them.
- Submit review records them. After that the sheet offers only Resubmit review: you can change your scores and resubmit until judging closes or the results are published, but a submitted review never goes back to being a draft.
- The comment goes to the organizers. Other judges never see it.
What you can and can’t see
You see your own assignments and your own scores, nothing else. Another judge’s review isn’t hidden on a page: it’s not reachable at all. The server answers “not found” for it, whatever URL you type or tool you use.
Conflicts
If you know the team, or have any other reason you shouldn’t judge a project, open it before you submit a review and choose I have a conflict, giving the reason. You’re taken off it, the organizers see it needs another judge, and it’s recorded in the audit log with your reason. You can’t be assigned a project from a team you’re on in the first place.
How your scores are used
Everyone reads a rubric a little differently: some judges are generous, some harsh, some use the whole scale and some stay near the middle. The calibration learns your habits from the projects you share with other judges and takes them out, so it doesn’t matter whether your 4 is someone else’s 3. If you give every project the same score, your scores carry no weight: they say nothing about which project is better. The method explains it all.
Assignment, scoring and calibration
This chapter is the repository’s
JUDGING.md,
included here as it stands so the two can’t drift apart. It covers how
reviews are handed out, how a scoresheet becomes one number, the model that
takes each judge’s habits out of the ranking, what that model guarantees and
where it stops. Every figure in it comes from
manage.py normalization_proof, which you can run yourself. If you want to
know what the calibration page shows you rather than how it’s computed, read
Reading the calibration page next.
How ballotbench hands out reviews, turns rubric scores into a ranking, and
why I think the ranking can be defended to a team that didn’t win. Every
number on this page comes from manage.py normalization_proof, which
recomputes it from the database; the output is in section 8.
1. Assigning reviews
portal/assignment.py is a pure function over plain records, so it’s tested
on thousands of random events without a database (tests/test_assignment.py).
- The idlest judge picks next. Of the judges who can still help, the one with the fewest reviews takes a project. That keeps loads within one review of each other wherever eligibility allows.
- They take the neediest project they may review. Fewest reviews so far, then fewest judges left who could take it, so scarce judges go where only they can help.
- Eligibility is enforced in the planner and again by a database trigger: never a project from your own team, only your tracks (a judge with no tracks takes any), never a project you already hold or stepped aside from.
- Repair. Greedy can corner itself: at the end, the idle judges may already hold the last projects that need someone. A repair pass moves one of the plan’s new reviews from the busiest judge to the idlest eligible one until no move narrows the gap. Existing reviews never move.
- Connectivity. Calibration compares judges through the projects they share. If the judge-project graph splits into groups that share nothing, their scales can’t be compared, so the planner adds bridging reviews (union-find over the graph) and the page says how many it added.
The organizer previews the plan, then applies it; on apply it’s recomputed from the database, never taken from the form. Manual changes are allowed, but a review that has been started can’t be taken back, since that would quietly delete a judge’s scores. A judge who spots a conflict steps aside; the project shows up as needing a top-up.
On the fixture, the plan for k = 3 is exactly the eight top-ups the fixture’s
two unfinished batches call for: prj_10, prj_15, prj_18, prj_19, prj_24, prj_29, prj_39, prj_40 each have two reviews. The progress dashboard flags
them as low coverage until the new reviews come in.
2. A review’s score
Each criterion has a weight and a range (1 to 5 by default). A review’s score is the weighted mean of its criteria, each first mapped to 0..1 by its own range, so a 1-10 criterion and a 1-5 criterion count by their weights, not by the width of their scales:
score = Σ_c w_c · (x_c − min_c) / (max_c − min_c) / Σ_c w_c
A review missing a criterion has no score and isn’t used. The fixture’s scale isn’t stated; its values run 2 to 5, so I assume 1 to 5.
The rubric freezes at the first score. A trigger refuses any change to criteria, weights or ranges once the event has a score. Weights tuned after reading the scores are a way to pick a winner.
3. Why not just average, or z-score
- Raw means reward drawing a lenient judge. With three reviews per project, one generous judge moves a project a long way.
- Per-judge z-scores fix leniency but assume every judge saw an average
batch. A judge who only saw the five best projects gets the best of those
pushed down to the middle. They also divide by zero for a judge who gives
the same score every time, which the fixture has (
jdg_07).
4. The model
Each review’s score y is modelled as
y_ij = a_j + s_j · q_i + e_ij, e_ij ~ N(0, v_j), q_i ~ N(0, 1)
q_i: project i’s quality, the thing we want.a_j: judge j’s leniency (offset).s_j: how strongly their scores follow quality (scale). A judge who separates good from weak sharply has a larges_j.v_j: how noisy they are.
The fit (portal/calibration.py) alternates two least-squares steps until q
stops moving (tolerance 1e-12):
q_i = Σ_j s_j (y_ij − a_j) / v_j / (1 + Σ_j s_j² / v_j) then centre and scale q to mean 0, sd 1
s_j = S_xy / (S_xx + 1), a_j = ȳ_j − s_j · q̄_j ridge least squares on the judge's own projects, s_j ≥ 0
v_j = (3 · var(y_j)/2 + RSS_j) / (3 + n_j) empirical-Bayes shrink of the noise
- The
q_istep is the posterior mean under the N(0, 1) prior: a project seen by few or noisy judges is pulled towards the middle instead of trusted blindly. That’s the answer to the unfinished batches. - A judge’s leniency is anchored by the projects they share with other judges, not by their own batch, which is what z-scores get wrong.
- The noise prior stops a judge with three reviews from being fitted exactly and then trusted infinitely.
- The ridge on
s_j(the+ 1) is there because my first version had none, and on the fixture it didn’t converge: a judge whose few projects landed close together in q got a huge scale, which dragged those projects closer, and so on. The penalty is in q’s units, which have no scale of their own, so it doesn’t break the invariances below.
This is the reviewer-calibration model used for NeurIPS reviewing (Lawrence, 2014; Ge, Welling & Ghahramani), fitted by alternating least squares rather than full Bayesian inference, which is plenty for tens of judges.
Judges it can’t learn from
These get s_j = 0. Because every term a judge contributes to q_i is
multiplied by s_j, a judge with s_j = 0 drops out exactly, as if their
reviews weren’t there. Raw means still include them; the calibration page
lists them with the reason.
| Flag | Meaning | On the fixture |
|---|---|---|
constant | every review the same score | jdg_07 (4 on every criterion, 3 reviews) |
single_review | one review: nothing to compare it with | jdg_01, jdg_23 |
discordant | their scores fall, or don’t rise, as everyone else’s rise on the projects they share | 9 judges |
What it guarantees (tested)
- Shift one judge (add a constant to all their scores): every
qis unchanged, to 1e-15. - Stretch one judge (multiply their scores by a positive factor): same.
- Shift and stretch every judge at once, each differently: same.
- Remove the constant judge: every
qis unchanged, exactly 0.
These hold because every step is equivariant: a judge’s a_j, s_j and
v_j absorb any affine change of their own scores, and the ridge and priors
are either in q’s units or built from the judge’s own spread.
5. How sure is the ranking?
The model’s own standard error, 1/√(1 + Σ s_j²/v_j), assumes each judge’s
offset and scale are known exactly. With three or four reviews per judge they
aren’t. So each run also bootstraps: resample each project’s reviews with
replacement, refit, 200 times, and record where each project lands. The pages
show the 90% range as “could be 4 to 12”. Where two projects’ ranges overlap,
the judges couldn’t really tell them apart, and the page says so rather than
printing a confident third decimal.
6. Does the data have a signal at all?
Each run also asks whether the judges agree about which projects are better more than chance would: it compares the spread of the project means with the spread after shuffling every score across the same judge-project slots (2000 shuffles). If they don’t, no method can produce a meaningful ranking, and the calibration page says so in plain words.
On the fixture they don’t (p ≈ 0.70). The fixture’s scores look like independent draws: projects’ means spread no more than shuffled scores do, and nine judges’ scores run against the consensus. So the honest result on the fixture is: the constant and single-review judges are handled, the duplicate is out, the ranking is computed, and almost every rank’s interval is wide (median 27 places). I’d rather show that than a crisp ranking that is noise.
7. The fixture’s traps
| Trap | What ballotbench does | Where you see it |
|---|---|---|
jdg_07 scores 4 on everything | flagged constant, weight exactly 0 | calibration page, judges.csv, the proof |
jdg_01 (and jdg_23) have one review | flagged single_review, weight 0 | same |
| two unfinished batches: 8 projects with 2 reviews | shrunk towards the middle, wider intervals, “low coverage” flag, 8 top-ups proposed. One of them, prj_24, was reviewed only by two judges the model ignores (both discordant), so it has nothing to be ranked by and is listed as left out until its top-up reviews arrive | progress dashboard, assignment preview, calibration page |
prj_41 repeats prj_07 (same team, title, repo, submitted later) | detected at import, duplicate_of set, out of the rankings until an organizer decides; its reviews still calibrate its judges | duplicates page, gallery badge, audit log |
| judge load 1 to 11 | noise shrinkage trusts busy judges more; new work goes to the idlest | judge table |
8. The proof on the fixture
docker compose exec web python manage.py normalization_proof, on a fresh seed:
Normalization proof: Sample Hack 2026, 126 reviews, 30 judges, 41 projects
Model: y = a_j + s_j*q_i + e, fitted in 210 iterations (converged), judge-project graph in 1 component(s).
Judges carrying no weight (s_j = 0):
jdg_07 3 review(s) constant
jdg_04 4 review(s) discordant
jdg_08 3 review(s) discordant
jdg_10 3 review(s) discordant
jdg_14 3 review(s) discordant
jdg_17 2 review(s) discordant
jdg_18 3 review(s) discordant
jdg_19 4 review(s) discordant
jdg_21 4 review(s) discordant
jdg_27 2 review(s) discordant
jdg_01 1 review(s) single_review
jdg_23 1 review(s) single_review
the other 18 judges: ok
Agreement: variance of project means 0.0080, 0.0089 on average when every score is shuffled across the same slots (permutation p = 0.698, 2000 shuffles).
The judges agree on which projects are better no more than chance would. Calibration takes out
judge habits; it can't create a signal the scores don't hold, so most rank moves below are noise.
Left out of the ranking: prj_41 (Dry Harbour) repeats prj_07; its 4 reviews still count towards calibrating their judges.
Left out of the ranking: prj_24 (Glass Beacon): all 2 of its reviewers carry no weight, so there's nothing to rank it by.
rank could be raw move project title n raw mean calibrated
1 1-39 31 +30 prj_07 Dry Harbour 5 0.583 0.832
2 1-33 5 +3 prj_37 Salt Loom 4 0.771 0.804
3 2-36 22 +19 prj_01 Glass Signal 3 0.611 0.774
4 2-37 9 +5 prj_08 North Drift 5 0.700 0.762
5 1-16 1 -4 prj_11 Salt Ledger 4 0.833 0.749
6 2-12 2 -4 prj_34 Iron Switch 3 0.833 0.746
7 3-36 15 +8 prj_09 Hollow Signal 3 0.639 0.741
8 4-36 25 +17 prj_12 Open Beacon 3 0.611 0.739
9 4-39 26 +17 prj_27 Flat Thread 3 0.611 0.732
10 2-29 4 -6 prj_25 Dry Relay 3 0.778 0.716
11 3-16 7 -4 prj_33 Slow Trail 3 0.750 0.709
12 4-31 8 -4 prj_21 Copper Kiln 3 0.722 0.708
13 6-34 18 +5 prj_02 Small Meadow 3 0.639 0.686
14 8-28 17 +3 prj_31 Salt Ferry 3 0.639 0.684
15 3-35 10 -5 prj_04 Green Switch 3 0.694 0.675
16 5-20 3 -13 prj_10 Still Beacon 2 0.792 0.666
17 3-38 21 +4 prj_35 Warm Beacon 5 0.617 0.652
18 3-36 30 +12 prj_29 Flat Relay 2 0.583 0.651
19 5-32 12 -7 prj_15 Copper Orbit 2 0.667 0.651
20 11-37 28 +8 prj_03 Deep Compass 3 0.583 0.634
21 1-29 6 -15 prj_16 Salt Kiln 3 0.750 0.624
22 11-29 14 -8 prj_36 Salt Drift 3 0.667 0.615
23 18-29 13 -10 prj_19 Small Relay 2 0.667 0.589
24 4-27 16 -8 prj_17 Small Loom 3 0.639 0.589
25 12-33 27 +2 prj_14 Green Lantern 5 0.600 0.582
26 5-36 19 -7 prj_18 Open Kiln 2 0.625 0.580
27 21-37 23 -4 prj_28 Flat Meadow 3 0.611 0.579
28 1-33 20 -8 prj_39 Paper Anchor 2 0.625 0.571
29 2-35 11 -18 prj_38 Deep Beacon 3 0.694 0.567
30 21-37 37 +7 prj_40 Slow Loom 2 0.500 0.564
31 4-38 36 +5 prj_30 Paper Harbour 3 0.528 0.559
32 18-37 34 +2 prj_06 Dry Compass 3 0.528 0.558
33 17-39 35 +2 prj_22 Dry Bridge 3 0.528 0.556
34 4-38 33 -1 prj_26 Amber Hours 3 0.556 0.544
35 5-38 24 -11 prj_32 Loud Ledger 3 0.611 0.529
36 12-39 29 -7 prj_13 Quiet Anchor 3 0.583 0.515
37 14-39 32 -5 prj_20 Paper Thread 3 0.556 0.504
38 28-39 38 prj_05 North Compass 3 0.472 0.500
39 24-39 39 prj_23 Slow Quarry 3 0.472 0.490
37 of 39 projects change rank. Kendall's tau, calibrated vs raw: 0.457
'could be' is a 90% bootstrap interval (each project's reviews resampled, 200 refits). Median width 27 places: on these scores most ranks are not distinguishable.
Invariance checks on these scores (max |change in q| over all projects):
PASS jdg_24 adds 0.3 to every score: 1.1e-15, ranking identical
PASS jdg_24 doubles every score: 0.0e+00, ranking identical
PASS every judge gets a random shift and stretch: 3.1e-15, ranking identical
PASS constant judge(s) jdg_07 removed: 0.0e+00, ranking identical
Kingmaker check: leaving out one judge at a time, 9 of 30 judges' absence would change who is in the top 3:
without jdg_11: prj_35 in, prj_01 out
without jdg_02: prj_16 in, prj_01 out
without jdg_20: prj_08 in, prj_01 out
without jdg_22: prj_08 in, prj_37 out
without jdg_26: prj_39 in, prj_07 out
without jdg_28: prj_08 in, prj_01 out
without jdg_03: prj_08 in, prj_01 out
without jdg_05: prj_08 in, prj_01 out
without jdg_06: prj_08 in, prj_01 out
Second opinion: the rubric read as pairwise picks (Bradley-Terry, portal/pairwise.py):
254 picks; no picks from jdg_01, jdg_07, jdg_23 (one review, or every pair tied)
Kendall's tau with the calibrated ranking 0.649, with raw means 0.630
PASS jdg_24's scores squared (not a shift or stretch): picks identical; the calibration's q moves by up to 0.04, since it only undoes linear habits
Benchmark on synthetic events with a known true order (40 projects, 12 judges, 3 reviews each,
one harsh judge who only sees the strongest projects). Kendall's tau against the truth:
calibration 0.830 raw means 0.677 per-judge z-scores 0.742 pairwise picks 0.799 (mean of 20 events)
calibration beats raw means in 20 of 20, z-scores in 20 of 20
Every invariance check holds.
The benchmark events are synthetic because the fixture has no ground truth. Each has 40 projects, 12 judges with their own leniency, scale and noise, and one harsh judge who only sees the eight strongest projects: the case that breaks z-scores. The model recovers the true order better than raw means and better than z-scores in every one of the 20 events, and on average better than the pairwise second opinion, which throws away how far apart a judge put two projects.
9. A second opinion: pairwise picks
The calibration undoes linear habits. A judge who squashes only the top of the scale is non-linear, and it can’t. So the calibration page shows a second ranking next to it that no way of using the scale can move.
Every judge who scored two projects differently has, in effect, picked the
better one. portal/pairwise.py collects those picks across all judges and
fits a Bradley-Terry model, P(A beats B) = p_A / (p_A + p_B), with Hunter’s
MM iteration and one virtual win and loss per project against a reference, a
weak prior that keeps an unbeaten project finite. Only the order of each
judge’s own scores enters, so any increasing transformation of any judge’s
scores leaves it unchanged; the proof squares one judge’s scores to show it.
The constant judge and the single-review judges make no picks at all.
What it gives up is magnitude: “A slightly better than B” and “A far better than B” are the same pick. That’s why it’s a second opinion rather than the ranking. Where the two agree, trust the rank more; where they disagree, the project’s rank depends on how you read the judges’ scales, and its “could be” range will usually be wide as well.
10. Does one judge decide the podium?
Before publishing, the calibration page runs the whole fit again once per judge, each time without that judge’s reviews, and lists every judge whose absence would change who is in the top three (or the top N, for N prizes). It takes about a second on the fixture.
This finds the case the other checks can’t: a judge the model trusts, whose
scores agree with everyone else’s on most projects, lifting one borderline
project onto the podium. A lone contrarian is not a kingmaker: a judge whose
scores run against the consensus is flagged discordant and carries no weight,
so leaving them out changes nothing (tests/test_calibration.py plants both
and checks each). A name on the list isn’t an accusation; it says the podium
rests on one person, which is where a second look, or one more review, is
worth it.
On the fixture, nine judges’ absence would each swap one podium place, consistent with scores that carry no real signal.
11. Spending the next reviews where they matter
Once a calibration has run, the assignment page lists the projects whose plausible rank range crosses the prize line (the number of prizes, or the top three): projects that could end up either side of it. Give each of these one more review adds exactly one review per contested project, from the idlest eligible judge. Spreading extra reviews evenly spends most of them on projects that can’t win and can’t miss; these are the ones where another opinion can change who wins. Run calibration again afterwards and the ranges narrow where it counted.
12. Explaining a rank to the team
After publication, each team can open How your project was scored: every
review of their project with the judge anonymized (“Judge 2”, shuffled per
project), what that judge gave them, what that judge gives a typical project
(their offset a_j), the difference, and the share of the final score that
review carried ((s_j²/v_j) / (1 + Σ s²/v)), with the pull towards the middle
as its own row. A judge who counted for nothing says why in words: gave every
project the same score, wrote one review, or scored against the consensus.
It shows weighted totals only, no per-criterion scores, comments or names,
and warns if scores changed after the published run.
13. Publishing results
Results are hidden from everyone but the event’s organizers until published, in the pages and the API (404 before then). Publishing freezes the result to one calibration run. Each run stores the SHA-256 of exactly the scores it read, in a canonical order, so anyone with the scores export can recompute it and check the published ranking came from those scores. Later runs change nothing public until someone publishes again, and every run and publication is in the audit log.
The published ranking is also available as one signed document,
/api/events/<slug>/results/signed: the event, the run, the scores’ digest
and every rank with its “could be” range, signed with the portal’s Ed25519
key. Anyone who saved a copy on results day can check it at /verify, or
offline with the public key at /.well-known/ballotbench-signing-key, and hold
the portal to it if the page ever says something else.
14. Known limits
- Linear judges only. The model corrects a judge who is lenient or who spreads scores widely. It can’t correct one who only compresses the top of the scale, or who uses the scale differently for different criteria (Wang & Shah, “Your 2 is my 1”).
- Collusion. Two judges who agree to push a project look like two judges who agree. Nothing statistical separates them; conflict-of-interest rules and the audit log are the defence.
- Thin data. With three reviews per project and a few per judge, ranks are uncertain. The intervals say so; they don’t fix it. More reviews per project do.
- One number per review. Calibration works on the weighted total. A per-criterion model would need far more reviews than a hackathon has.
References
- D. Hunter, “MM algorithms for generalized Bradley-Terry models”, Annals of Statistics, 2004.
- N. Lawrence, “Reviewer calibration for NIPS”, 2014. https://inverseprobability.com/2014/08/02/reviewer-calibration-for-nips
- H. Ge, M. Welling, Z. Ghahramani, “A Bayesian model for calibrating conference review scores”. https://mlg.eng.cam.ac.uk/hong/unpublished/nips-review-model.pdf
- M. Roos, J. Rothe, B. Scheuermann, “How to calibrate the scores of biased reviewers”, AAAI 2011.
- J. Wang, N. Shah, “Your 2 is my 1, your 3 is my 9”, 2018. https://arxiv.org/abs/1806.05085
Reading the calibration page
The calibration page is where an organizer turns a pile of reviews into a
ranking, and it’s the page I’d want open when a team asks why they came
fourth. It’s at Manage → Calibration and results, or
/events/<slug>/manage/calibration. Only the event’s organizers and site
admins can open it; anyone else gets a 403.
This chapter goes through it block by block, using the fixture event
(Sample Hack 2026) as it looks straight after docker compose up. The
numbers below are from a freshly seeded stack; yours will match until
someone changes a score.
Running it
The button at the top says Run calibration the first time and Run it again on the current scores after that. A run reads every submitted review that has a score for every criterion; drafts and half-filled scoresheets are ignored. It fits the model, bootstraps the rank intervals (200 refits) and runs the agreement test, which takes a few seconds on the fixture.
Every run is kept, and a run is never updated. The page always shows the
latest one. That matters for two reasons: you can run it as often as you
like while judging is still going, and a published result can’t change
underneath you, because publishing points at one run and later runs are new
rows. Each run also writes a calibration.run row to the audit log with its
fingerprint, the number of reviews, the agreement p-value and the judges it
flagged.
The run facts
The first block is a short list of facts about the run.
| Line | What it says | On the fixture |
|---|---|---|
| When | the time the run was made (UTC) and who made it | your run |
| Reviews read | submitted, complete reviews the run used | 126 |
| Fingerprint | SHA-256 of exactly the scores the run read | 397dd35f…178f4a |
| Judges linked | whether every judge can be compared with every other | Yes |
| Agreement | whether the judges agree more than chance would | no (p = 0.70) |
Fingerprint
The fingerprint is a SHA-256 over the scores the run read, written out in a
fixed order. It’s there so that a ranking can be checked against the scores
it claims to come from: if one score changes, the fingerprint changes. The
same fingerprint is stored with the run, shown on the public results page,
returned by the results API as input_digest, and written into every row of
results.csv.
On a freshly seeded fixture it’s always:
397dd35f376eb052b0dfcd8fda8309f46bad4771e9caa4e78035175807178f4a
The fingerprint doesn’t prove the scores are the ones the judges meant to
give. It proves the ranking came from these scores. For the first question,
the audit log records every review.save and review.submit with the
before and after values, and manage.py verify_audit checks nobody has
edited that log. Recompute the fingerprint
below shows how to check it by hand.
Judges linked
Calibration compares judges through the projects they share. If the judges split into groups that share no project, directly or through other judges, their scales can’t be put side by side: a lenient group and a harsh group look exactly like a strong batch and a weak one.
- Yes means every judge is connected. The fixture is.
- No is flagged, with the advice to hand out bridging reviews. Open Hand out reviews: the planner detects the split, adds bridging reviews until the groups connect, and says how many it added. Once those reviews are in, run calibration again.
Agreement
This line answers the question to ask before looking at any rank: do the judges agree about which projects are better any more than chance would? The test shuffles every score across the same judge-project slots 1000 times and counts how often a shuffle spreads the project means at least as far apart as the real scores do. That fraction is the p-value.
- p below 0.05: “The judges agree on which projects are better far more than chance would.” There’s a real signal. The ranking still has uncertainty, which the “could be” column shows.
- p of 0.05 or more: “The judges agree no more than chance would”, with a flag. The ranking is computed anyway, but most of its order is noise.
- No reviews yet if the run read nothing.
The fixture reads p = 0.70 (the page and normalization_proof run the same
test: 2000 shuffles, the same seed): its scores behave like independent draws.
When p is high, I would not publish a ranking to three decimals. What I’d do, roughly in this order:
- Treat overlapping ranks as ties. Read the “could be” column, not the rank. If you must name winners, name a group whose ranges sit clearly above the rest, and say it is a group.
- Get more reviews. The signal grows with reviews per project. Raise k in the event settings, hand out more reviews, then run it again.
- Look at the flagged judges (next section). Nine discordant judges out of thirty, as on the fixture, says the judges weren’t reading the rubric the same way, or the rubric doesn’t separate the projects.
- Say so when you announce. “The judges couldn’t separate projects 4 to 20” is a fair thing to tell teams, and the public results page already shows each project’s range.
Calibration can take out a judge’s leniency and scale. It can’t create agreement that isn’t in the scores, and the page says that in as many words.
Judges the model couldn’t learn from
This table lists judges whose reviews carry no weight in the calibrated scores: their fitted scale is exactly 0, which removes them from every project’s score as if they hadn’t reviewed. Their reviews still count in the raw means and in each project’s review count. Each row shows the judge’s name (or email), how many reviews they wrote, and why.
| Flag | The page says | What it means | What to do |
|---|---|---|---|
| constant | Gave every project the same score, so their scores say nothing about which is better. | every review has the same weighted score | ask them; if they misread the rubric, have them rescore, otherwise top up their projects with another judge |
| single review | Only one review: there’s nothing to compare it with. | one review can’t reveal a judge’s leniency or scale | give them more reviews that overlap with other judges |
| discordant | Their scores don’t rise with everyone else’s on the projects they share. | their fitted scale came out at zero or below | read their comments first; more overlapping reviews settle it either way |
On the fixture, the page lists twelve judges. It shows names; the fixture ids
are here so you can match them to fixtures.json and to the method chapter:
| Fixture id | Name on the page | Reviews | Flag |
|---|---|---|---|
jdg_07 | Iva Petrova | 3 | constant |
jdg_01 | Tomas Varga | 1 | single review |
jdg_23 | Anya Sokolova | 1 | single review |
jdg_04 | Noor Haddad | 4 | discordant |
jdg_08 | Marek Nowak | 3 | discordant |
jdg_10 | Hiro Tanaka | 3 | discordant |
jdg_14 | Emeka Adeyemi | 3 | discordant |
jdg_17 | Bruno Costa | 2 | discordant |
jdg_18 | Lars Berg | 3 | discordant |
jdg_19 | Mira Kaur | 4 | discordant |
jdg_21 | Sana Aziz | 4 | discordant |
jdg_27 | Leila Nasser | 2 | discordant |
jdg_07 gave 4 on every criterion of all three projects. A judge who
gives everyone the same score tells you nothing about which project is
better, so the model gives them no weight (a per-judge z-score would divide
by zero here). The side effect is on the projects they reviewed: one of
them, prj_19 (Small Relay), has only two reviews, so its calibrated score
rests on the other judge alone.
jdg_01 and jdg_23 wrote one review each. With one review there is
nothing to separate a judge’s leniency from the project’s quality, so they
get no weight either. jdg_01’s one review is of prj_07, which matters in
the first worked example below.
Discordant is not an accusation. It means that on the two to four projects a judge shared with others, their scores went down, or stayed flat, while everyone else’s went up. With so few reviews, one honest disagreement can do that. On the fixture it’s mostly noise, which fits the agreement line. The flags are recomputed on every run, so a judge can move between ok and discordant as reviews come in.
Each judge’s fitted numbers (offset, scale, noise and flag) are in the
Judges export, judges.csv.
The ranking
Below the judges is the ranking itself, one row per project that has at least one review in the run. Projects with no reviews at all don’t appear.
| Column | What it is |
|---|---|
| Rank | the calibrated rank among ranked projects; ties, which are rare, go to the higher raw mean |
| Could be | the 90% bootstrap range for the rank, “a to b” |
| Raw rank | the rank by plain mean of weighted scores, with “up n” or “down n” |
| Project | the title, linked to the project page, and any flag |
| Reviews | submitted, complete reviews of this project in the run, flagged judges included |
| Raw mean | the plain mean of those reviews’ weighted scores, 0 to 1 |
| Calibrated | the calibrated score, mapped back onto the rubric’s 0 to 1 scale |
Rank and calibrated score
The calibrated score is the model’s estimate of each project’s quality with every judge’s leniency and scale taken out. It’s put back on the 0 to 1 scale (the average project, plus so many standard deviations of the project means) so it reads like a rubric score, but it’s only comparable within one run. A project seen by few judges, or by noisy ones, is pulled towards the middle rather than trusted on thin evidence.
Could be
This is the column I’d read first. For each run, every project’s reviews are resampled with replacement and the whole model refitted, 200 times, and the column shows the range the project’s rank falls in 90% of the time. A narrow range means the scores place the project consistently; a wide one means they can’t tell it from its neighbours. Where two projects’ ranges overlap, the judges couldn’t really separate them.
On the fixture, most ranges are wide (half of them span 27 places or more), which is the agreement line again, seen project by project.
One caution: the bootstrap can only resample the reviews that exist. A project with two reviews has only three possible resamples, so a narrow range on a low-coverage project is less reassuring than it looks.
Raw rank, up and down
The raw rank orders the same projects by their plain mean. The label next to it compares the two: “up 30” means the project is 30 places higher after calibration than on raw means, “down 4” means 4 places lower, and no label means it didn’t move. A big move with a narrow “could be” is calibration doing its job: a project that drew a harsh judge, say. A big move with a wide “could be” is the model reshuffling noise.
Low coverage
A project with fewer reviews in the run than the event’s k (reviews per project, 3 by default) is flagged low coverage. On the fixture, eight projects from the two unfinished batches have two reviews each:
| Project | Title | Reviewers | Rank | Could be | Raw rank |
|---|---|---|---|---|---|
prj_10 | Still Beacon | jdg_15, jdg_29 | 16 | 5 to 20 | 3, down 13 |
prj_15 | Copper Orbit | jdg_13, jdg_10 (discordant) | 19 | 5 to 32 | 12, down 7 |
prj_18 | Open Kiln | jdg_24, jdg_04 (discordant) | 26 | 5 to 36 | 19, down 7 |
prj_19 | Small Relay | jdg_29, jdg_07 (constant) | 23 | 18 to 29 | 13, down 10 |
prj_24 | Glass Beacon | jdg_18, jdg_19 (both discordant) | left out | ||
prj_29 | Flat Relay | jdg_24, jdg_09 | 18 | 3 to 36 | 30, up 12 |
prj_39 | Paper Anchor | jdg_26, jdg_18 (discordant) | 28 | 1 to 33 | 20, down 8 |
prj_40 | Slow Loom | jdg_26, jdg_24 | 30 | 21 to 37 | 37, up 7 |
Most of them moved towards the middle. That is the shrinkage working: two
reviews, several of them from judges who carry no weight, aren’t enough
evidence to keep a project near the top or the bottom. prj_10 had the
third-best raw mean on two reviews and ends up 16th, with a range of 5 to 20.
To clear the flag, open Hand out reviews. With k = 3, the plan for the fixture is exactly one top-up for each of the eight. The flag stays until the new reviews are submitted and you run calibration again.
Left out
Some projects are in the table but not ranked. They’re at the bottom, greyed out, with left out: and a reason. Their reviews still help calibrate the judges who wrote them; the project just doesn’t get a rank.
| Reason | When |
|---|---|
| duplicate | the project is flagged as a duplicate of an earlier one and nobody has put it back |
| not submitted | the project has reviews but is a draft |
| no usable reviews | every one of its reviews is from a judge who carries no weight |
On the fixture there are two:
prj_41Dry Harbour, left out: duplicate. Team CopperLedger submitted Dry Harbour twice, with the same title and repository:prj_07on 1 March at 04:29 andprj_41three minutes before the deadline. The import flags the later one. Its four reviews still count towards calibrating their judges. Open Duplicates to confirm it or put it back; if you put it back, run calibration again and it’s ranked like any other project.prj_24Glass Beacon, left out: no usable reviews. Both of its reviews are from discordant judges (jdg_18andjdg_19), so the model has nothing to go on. The calibrated score it shows is just the average project, and means nothing. It needs a review from another judge.
Two worked examples
prj_07 Dry Harbour is ranked first, and I wouldn’t announce it. Its raw
rank is 31, so calibration moved it up 30 places. It has five reviews, but
two are from discordant judges (jdg_19, jdg_21) and one is jdg_01’s
single review, so its calibrated score rests on two judges, jdg_12 and
jdg_26, who both scored it well relative to how they scored everything
else. Its “could be” is 1 to 39: nearly the whole field. On this data, first
place means “the two judges who count liked it”, not “it won”.
prj_11 Salt Ledger and prj_34 Iron Switch are the ones I’d trust.
They had the best raw means (raw rank 1 and 2) and calibration moved each of
them down four places, to 5 and 6. But their ranges are among the narrowest
on the page: 1 to 16 and 2 to 12. Whatever the resampling does, the data
keeps putting them near the top. If I had to name a shortlist from the
fixture, I would build it from the ranges, not the ranks.
Publishing
Until results are published, the results page and the results API answer 404 to everyone except the event’s organizers, who see the latest run with a “Preview” banner.
Publish results freezes the public results to the run on this page:
- It records the time (on the database’s clock) and the run on the event,
and writes a
results.publishaudit row with the run and its fingerprint. /events/<slug>/resultsbecomes public. It shows each ranked project’s rank, title, team, track, calibrated score, number of reviews, the range it could have landed in, and the fingerprint. Left-out projects aren’t listed. Judge names and flags are never public./api/events/<slug>/resultsbecomes public too, with the same projects plus each one’s raw rank, raw mean and standard error (see the API).- Running calibration again afterwards changes nothing public. The page then offers Publish run n instead, so moving the public result to a newer run is always a deliberate act, and it’s audited.
- Hide the results again takes them down (
results.unpublishin the audit log). Both pages go back to 404 for everyone but organizers.
One thing to watch: the Results and Judges exports always describe
the latest run, not the published one. Their run column says which run a
file came from.
Before I press publish I check, in order: judges linked, no low-coverage flags left, duplicates decided, the flagged judges looked at, and what the agreement line says. The last one decides how I word the announcement.
Recompute the fingerprint
Anyone with the scores export can check a fingerprint. The export is organizer only, so in practice an organizer downloads it and hands it to whoever wants to check: a team, another organizer, a hackathon judge.
This is what a run hashes, from run_calibration in
src/portal/results.py
and digest in
src/portal/calibration.py:
- one row per submitted review with every criterion scored, as
[review_id, judge_email, project_id, [[criterion_key, value], ...]], the pairs sorted by criterion key; - the rows sorted, which puts them in
review_idorder; - serialized with
json.dumps(rows, separators=(",", ":"))and hashed with SHA-256.
Every piece of that is a column in scores.csv: review_id, judge_email,
project_id (the portal’s numeric id, not the fixture’s prj_ id), and one
column per criterion, headed by its key. A review the run skipped has an
empty submitted_at or an empty weighted_0_1.
curl -s -H "Authorization: Bearer bb_demo_organizer_5c1e0a" \
-o scores.csv http://localhost:8080/api/events/sample-hack-2026/export/scores.csv
python3 fingerprint.py scores.csv
# fingerprint.py: recompute a calibration run's input digest from scores.csv
import csv, hashlib, json, sys
FIXED = {"review_id", "judge_email", "project_id", "project", "submitted_at", "weighted_0_1", "comment"}
with open(sys.argv[1], newline="", encoding="utf-8") as f:
reader = csv.DictReader(f)
criteria = [c for c in reader.fieldnames if c not in FIXED]
rows = []
for r in reader:
if not r["submitted_at"] or not r["weighted_0_1"]:
continue # drafts and incomplete reviews aren't read
rows.append([int(r["review_id"]), r["judge_email"], int(r["project_id"]),
sorted([c, int(r[c])] for c in criteria)])
canon = json.dumps(sorted(rows), separators=(",", ":"))
print(hashlib.sha256(canon.encode()).hexdigest())
On a freshly seeded stack this prints 397dd35f…178f4a, the fixture’s
fingerprint. If it doesn’t match a run’s fingerprint:
- A score changed after the run. That’s what the fingerprint is for. The audit log says which review, when, and who. The fixture’s judging window has no close date, so its judges can still change scores.
- A judge’s email changed. The export uses the current address.
- A cell was escaped. The export puts a quote in front of any cell that
starts with
=,+,-or@, so a spreadsheet won’t run it as a formula. Criterion keys and emails normally never start with those; if one does, strip the leading'first.
Compare with the fingerprint of the run you care about: the published one is
input_digest in the results API, the latest one is on this page and in
results.csv.
Architecture
ballotbench is one Django application in front of one Postgres database,
started by one docker compose up. Pages are rendered on the server; the
JSON API sits next to them and answers the same questions with the same
rules. There is no JavaScript framework, no queue beyond a webhook outbox
table, no cache and no outside
service, because a hackathon portal has hundreds of users, not millions, and
every moving part is something a volunteer organizer has to keep running.
browser ──► gunicorn :8080 ──► Django ──► Postgres 18 ◄── deliver_webhooks ──► your receivers
curl ─┘ (web container) │ (db container) (webhooks container, same image)
├─ pages portal/views, participant, judge, organizer, progress, results,
│ oversight, voting, comments, explain, records, embed, webhooks
├─ API portal/api, judge, results, voting, comments, records, bundles, exports
└─ rules portal/access + triggers in the database
Components
| Module | What it does |
|---|---|
models.py | the schema, with its unique and check constraints |
migrations/0002-0005, 0012 | the triggers (12 in all): deadline, team rules, judging rules, the audit chain, and the voting rules in 0012 |
access.py | roles, the scoped querysets every view starts from, the status-code rules, the deadline check on the database clock |
auth.py | bearer tokens (hashed), for the API and, via a middleware, the pages |
audit.py | writes one audit row per change, in the change’s transaction |
importer.py, commands/seed.py | the idempotent fixture import and the demo accounts |
participant.py | teams, invite links, the project form |
judge.py | the judge’s queue, the scoresheet, recusal, the judging API |
organizer.py | event settings, tracks, prizes, rubric, role invites |
assignment.py | the review planner, a pure function (see JUDGING.md) |
progress.py | applying plans, manual changes, the progress dashboard |
calibration.py | the calibration model, bootstrap intervals, the agreement test, a pure module |
results.py | calibration runs, publishing, the results pages and API |
oversight.py | audit log page, duplicate resolution, the exports page |
exports.py | CSV exports |
duplicates.py | duplicate submission detection |
voting.py, ratelimit.py, mail.py | the community vote, rate limits counted in Postgres, the offline mail outbox |
comments.py | project comments and their moderation |
pairwise.py, explain.py | the Bradley-Terry second opinion; a team’s own “how we were scored” page |
records.py, signing.py, embed.py, bundles.py, webhooks.py, tokens.py | T4: signed records and certificates, the embeddable gallery, event bundles, webhooks, personal API tokens |
The two pieces with real logic, the planner and the calibration, take plain tuples and return plain objects. They know nothing about Django, so they’re tested on thousands of random inputs in milliseconds, and the database layer around them is thin.
A request, end to end
GET /api/judge/scores?judge=jdg_24 with judge B’s token:
- gunicorn hands the request to Django.
ATOMIC_REQUESTSopens a transaction for the whole request. - DRF runs
BearerTokenAuthentication: SHA-256 the token, look up the hash, refuse a revoked token or an inactive user (401). api.judge_scoresresolvesjdg_24to a user. It isn’t the caller, and the caller organizes no event, so it raisesPermissionDenied: 403. A name that matches nobody also gets 403, so the answer doesn’t reveal which judges exist.- Without
?judge=, the queryset starts fromaccess.judge_assignments(user): the caller’s own assignments, in events where they still hold a judge membership. Reviews are filtered from that, never looked up by id.
POST /api/events/sample-hack-2026/projects with the participant’s token:
- Authenticated as above.
- The caller needs a team in the event, or it’s 403.
access.submissions_closed_reasonasks the database fornow()against the event’s window. Past the close: 409 with the close time.- If it were open, the write happens inside
access.guarded(), a savepoint that turns a trigger’s refusal into a 409 or 422. The deadline trigger checks the same clock again, so a request that races the deadline by a millisecond is still refused. - An audit row is written in the same transaction.
Where each rule is enforced, and why there
| Rule | App (for the message) | Database (for the guarantee) |
|---|---|---|
| deadline | submissions_closed_reason, on the DB clock | project_deadline trigger |
| a judge sees only their own work | querysets start from judge_assignments(user); naming another judge is 403 | assignments can’t exist for non-judges or own-team projects |
| one team per person, team size | form checks | unique constraint, team_member_rules trigger with a row lock |
| no judging your own team | planner skips it | assignment_rules, team_member_zz_not_judge, membership_judge_not_member |
| scores in range | form and API validation | score_rules trigger |
| rubric frozen once scored | page hides the inputs | rubric_frozen trigger |
| audit log can’t change | no update code exists | audit_readonly, audit_no_truncate, the hash chain |
| results hidden until published, and while voting is open | results.visible_run, voting.public_tallies; results.publish refuses while voting is open | (read rule; no write to guard) |
| a vote: window, budget, own team, eligible project, confirmed and not voided | voting.cast | vote_rules trigger, locking the voter row |
| one ballot per account and per inbox | voting.normalize_email, request_link | unique constraints on portal_voter |
| rate limits | ratelimit.allow, counted in portal_ratehit | (shared across workers because it is in the database) |
The database is the last word because a portal has more write paths than anyone remembers: pages, the API, the admin, a management command, a shell at 2 a.m. during the event. The app’s checks are for clear error messages; the triggers are for when someone forgets.
Status codes
The same everywhere, pages and API:
- 400: the request is malformed: a required field missing, a wrong type, a track from another event.
- 401: no credentials, or a bad token, on a route that needs them. Pages redirect a browser to the login page instead.
- 403: you’re signed in but your role can’t do this, or you named another
person’s data (
?judge=). - 404: the object exists but isn’t yours to see: another judge’s assignment, another team’s draft, unpublished results. Saying 403 would confirm it exists.
- 409: the window for this action is closed, or a conflict of interest.
- 422: values out of range.
Trade-offs
- Server-rendered pages, not a SPA. Fewer moving parts, works without JavaScript, one place for the rules. The cost is less interactivity, which a judging form doesn’t need.
- Triggers in SQL. They’re less familiar to many Django developers than model methods, and they tie us to Postgres. In exchange, the rules hold for every write path, including ones that don’t exist yet. Each trigger’s migration explains it in a docstring.
- The calibration in pure Python, no numpy. A hackathon has tens of judges and hundreds of reviews: a fit takes 16 ms and 200 bootstrap refits take about 3 seconds. One dependency fewer in the image.
- No email leaves the box. The portal must run offline, so invite links
are shown once to the person who makes them, to send however they like.
The one thing that has to be mailed, a voter’s confirmation link, lands in
an outbox table that site admins read in the admin. A real deployment sets
DJANGO_EMAIL_BACKENDto SMTP. - Sessions and tokens side by side. Browsers use sessions (with CSRF); scripts and the checker use bearer tokens (no cookies, so no CSRF risk). Pages accept tokens too so that the isolation probe tests what a browser gets.
- Two processes. gunicorn with three workers, and the
webhookssender (same image). Nothing else runs in the background: calibration takes a few seconds and runs inside the organizer’s request.
Beyond T2 (tier T4)
| Module | What it does |
|---|---|
signing.py | the Ed25519 key (one file, made on first use with mode 0600), canonical JSON, sign and verify; a pure module |
records.py | signed participation records for judges and participants, /verify, the public key, certificates |
embed.py | the frameable gallery /embed/<slug> and /embed.js |
bundles.py, commands/export_event, commands/import_event | whole-event bundles out, and in as a new event through importer.py |
webhooks.py, commands/deliver_webhooks | the outbox sender, the address checks, the organizer’s page |
- Records are checkable without us. What’s signed is the record’s
canonical JSON (sorted keys, no spaces, UTF-8), so the Python snippet on
/verifychecks one with nothing but the public key. A record never holds a score: judges’ scores stay private even from the people they judged. - Framing is opt-in per view.
X_FRAME_OPTIONS = "DENY"stays the default; only the two embed views are exempt, and the embed renders as an anonymous visitor whatever cookies arrive, so it can’t leak a draft into someone else’s page. The iframe reports its height withpostMessage, andembed.jsaccepts it only from that iframe and the portal’s origin. - Webhooks are the one background worker.
audit.recordqueues aWebhookDeliverynext to the audit row, in the same transaction, so the outbox and the log can’t disagree. A separatewebhooksservice (same image) sends them: it claims a round, one due delivery per webhook, by leasing them underSKIP LOCKEDrow locks (so two senders never claim the same one), commits, and sends them all at once with nothing locked. The web process never makes an outbound request. - SSRF. A webhook URL must resolve only to public addresses, when it’s saved and again at every send, and the connection goes to the address that was checked (no second DNS lookup to rebind), with one 10 s deadline for the whole attempt and no redirects.
Data model
Postgres 18. Every rule that must hold whatever code path writes a row lives
in the database: foreign keys, unique and check constraints, and triggers.
The application checks the same rules first so it can answer with a clear
409 or 422, but it doesn’t have to be right for the data to stay right.
Models are in src/portal/models.py; triggers are in the migrations
0002 to 0005, plus the voting rules in 0012 (12 triggers in all).
Tables
People and access
portal_user: an account. Signs in with email (unique, and a check
constraint keeps it lowercase). is_staff is the global admin role; every
other role is per event. email_confirmed_at is set when the person proves
they read the inbox (a signed confirmation link, or a completed password
reset); account-mode community votes need it.
portal_apitoken: user, label, token_hash (unique), created_at,
revoked_at. Only the SHA-256 of a token is stored; the token itself is shown
once.
portal_membership: user, event, role (participant, judge or
organizer), tracks (many-to-many; for judges, the tracks they may review,
empty meaning any), external_id.
- unique
(user, event, role) - unique
(event, role, external_id): fixture judge ids likejdg_24 - trigger
membership_judge_not_member: nobody becomes a judge of an event they’re on a team in
portal_roleinvite: single-use link that grants judge or organizer in one
event. token_hash (unique), expires_at, used_by, used_at, tracks.
Check: used_by set only with used_at.
Events
portal_event: slug (unique), name, external_id (unique),
submissions_open/close, judging_open/close, voting_open/close,
results_published_at, published_run (the calibration run the public
results are frozen to), voting_mode (off, account or email),
vote_credits (each voter’s quadratic budget), reviews_per_project (k),
max_team_size.
- checks: each window opens before it closes; k ≥ 1; team size ≥ 1
portal_track: event, name, external_id. Unique (event, name)
and (event, external_id).
portal_prize: event, optional track, name, description. Unique
(event, name).
Teams and projects
portal_team: event, name, external_id. Unique
(event, external_id). Names are not unique: the fixture has three
different teams called StillTrail.
portal_teammember: team, event, user, joined_at.
- unique
(event, user): one team per person per event - trigger
team_member_rules: copies the team’s event intoevent(so the unique constraint means what it says), locks the team row, and refuses a member pastmax_team_size. The lock makes two simultaneous invite acceptances queue rather than both squeeze in. - trigger
team_member_zz_not_judge: a judge of the event can’t join a team in it
portal_teaminvite: team, token_hash (unique), created_by,
expires_at, used_by, used_at. Single use, 72 hours or until submissions
close, whichever is first.
portal_project: event, team, track, title, tagline,
summary, description, repo_url, demo_url, tags (text array),
status (draft or submitted), submitted_at, duplicate_of (self),
duplicate_cleared (an organizer said it isn’t one; never flagged again),
external_id.
- unique
(event, external_id) - check: submitted ⇔
submitted_atset - check: not its own duplicate
- index
(event, status)for the gallery - trigger
project_same_event: team and track belong to the project’s event - trigger
project_deadline: see below
Judging
portal_rubriccriterion: event, key, name, weight,
min_value, max_value, position.
- unique
(event, key); checks: weight > 0, min < max - trigger
rubric_frozen: once any score exists in the event, no criterion can be added, deleted, reweighted or re-ranged. Renaming is allowed.
portal_judgeassignment: event, judge, project, status
(pending, done, recused), source (seed, auto, manual).
- unique
(judge, project) - trigger
assignment_rules: the project is in the event, the judge has a judge membership in it, and the judge isn’t on the project’s team
portal_review: one per assignment (assignment unique). comment,
submitted_at (null while a draft).
portal_score: review, criterion, value.
- unique
(review, criterion) - trigger
score_rules: the criterion belongs to the review’s event and the value is inside its range criterionisON DELETE RESTRICT: a scored criterion can’t be deleted alone, but an event can be deleted whole
Calibration
portal_calibrationrun: event, created_at, created_by, method,
params (JSON: priors and the rubric weights used), input_digest (SHA-256
of exactly the scores read), connected, converged, signal_p,
n_reviews.
portal_calibratedproject: run, project, n_reviews, raw_mean,
quality (q), display (q on the 0..1 scale), se, rank, raw_rank,
rank_low, rank_high (90% bootstrap interval), excluded (why it isn’t
ranked: duplicate, not submitted, no usable reviews). Unique (run, project).
portal_judgecalibration: run, judge, n_reviews, offset,
scale, noise, flag (ok, constant, single_review, discordant). Unique
(run, judge).
A run is never updated. Publishing points event.published_run at one.
Community voting and comments
portal_voter: one ballot in one event. event, user (account mode)
or email as typed plus email_normalized (email mode), token_hash (the
SHA-256 of the one-time confirmation link), confirmed_at, ip,
user_agent, voided_at, voided_reason.
- unique
(event, user)and(event, email_normalized): one ballot per account and per inbox - checks: an account or an address; a voided ballot has a reason
portal_vote: voter, project, votes (≥ 1; a zero is no row).
Unique (voter, project).
- trigger
vote_rules(migration 0012): locks the voter row, then refuses the write outside the voting window (BB409), for a voided or unconfirmed voter (BB409), for a project that isn’t a submitted, non-duplicate project of the voter’s event (BB422), for the voter’s own team (BB423), or when the sum of votes² would passvote_credits(BB409).
portal_comment: project, author, body, created_at, hidden_at,
hidden_by.
- checks: body not empty and at most 2000 characters
portal_ratehit: key, created_at. One counted action for a rate
limit; ratelimit.allow() deletes a key’s rows once they leave its window.
portal_outboundemail: to, subject, body, created_at. Mail the
portal would send, kept for site admins to read (portal.mail.OutboxBackend).
The audit log
portal_auditlog: seq, ts, actor, event, action,
object_type, object_id, before and after (JSON), ip, prev_hash,
row_hash.
- trigger
audit_append(BEFORE INSERT): takes a transaction-scoped advisory lock, setsseqto the next number,tsto the clock,prev_hashto the previous row’s hash, androw_hash = sha256(prev_hash | seq | ts | actor | event | action | object | before | after | ip). - triggers
audit_readonlyandaudit_no_truncate: UPDATE, DELETE and TRUNCATE are refused. actorandeventhave no foreign-key constraint on purpose: deleting a user or an event must not delete, or be blocked by, its history.manage.py verify_auditrecomputes the chain and names the first row that doesn’t fit.
Why seq and not the primary key: ids are handed out by a sequence before the
trigger takes its lock, so two concurrent inserts can commit in the opposite
order to their ids. seq is assigned under the lock.
The deadline, in the database
project_deadline runs before every INSERT, UPDATE and DELETE on
portal_project. If the database clock (now()) is past the event’s
submissions_close, it refuses a new project, a deleted one, or any change to
content (title, text, links, tags, track, team, status). It allows changes to
duplicate_of and duplicate_cleared, which are organizer bookkeeping.
The one bypass is the session setting ballotbench.import, which only the
fixture importer and delete_event set, with SET LOCAL so it ends with
their transaction, and both write an audit row saying so.
Refusals and what the API returns
Triggers raise with their own SQLSTATE, which access.guarded() maps to a
status code:
| SQLSTATE | Meaning | HTTP |
|---|---|---|
BB409 | a window is closed (deadline, frozen rubric) | 409 |
BB410 | team is full | 409 |
BB423 | conflict of interest | 409 |
BB422 | an inconsistent reference (wrong event, out-of-range score) | 422 |
BB403 | the audit log is append only | 403 |
Getting data in
- The fixture:
manage.py seedon every boot. Imported rows keep their fixture ids inexternal_id, withUNIQUE (event, external_id), so a second run inserts nothing and never overwrites what people changed. Fixture projects come in as submitted, through the import bypass, because they were submitted before a close date that has passed. - Another event in the fixture’s shape:
portal.importer.import_event(data)is the same importer;manage.py seed --fixtures path.jsonloads one.
Getting data out
- CSV, organizer only, audited, at every stage: registrations, teams,
projects, assignments, scores (each criterion and the weighted score),
results (raw and calibrated, rank intervals, the run’s digest), judges
(per-judge calibration), audit (with the hash chain).
GET /api/events/<slug>/export/<kind>.csv. - JSON: everything the pages show is in the API (
/api/docs). - The database itself: plain Postgres,
pg_dumpworks.
Beyond T2 (tier T4)
portal_webhook: event, url, secret (64 hex characters; it keys
the HMAC, so it’s stored as is, shown once and never written to the audit
log), actions (text array of action prefixes, empty meaning every change),
active, created_by (who it sends on behalf of: it pauses once they no
longer organize the event, and whoever resumes it takes it over), created_at.
portal_webhookdelivery: the outbox. webhook, action, payload
(JSON: the audit row’s action, event, actor, object, before and after),
status (pending, delivered, failed), attempts, next_attempt_at
(defaults to the database clock; while a sender has it claimed, the end of
its lease), last_error, created_at. Index
(status, next_attempt_at) for the sender. A row is written by
audit.record in the change’s transaction, so it rolls back with it. Failed
sends wait 30 s, then twice as long each time; the eighth failure is final.
Signed records aren’t stored: each is built from the tables above when
asked for, and the issuance is an audit row (record.issue). The signing key
is a file (BALLOTBENCH_SIGNING_KEY_FILE, /data/signing_key.pem in
Docker), not a table, so a database dump doesn’t carry it.
Event bundles (bundles.py) are the fixture’s shape plus format,
prizes, criteria (with weights and ranges), assignments (those without
a review) and, per row, everything the fixture leaves out: event dates,
judges’ tracks, every project field with status and duplicate_of, and
each review’s submitted flag, time, status and comment. Rows are named by
external_id, or prj_<pk> and the like. Importing one as a new event goes
through import_event(..., new=True) with the deadline bypass, in one
transaction; URLs, emails and every enum are validated first, and the bypass
is switched off again when the import ends rather than when the request does.
Threat model
Who attacks a hackathon portal, how, and what ballotbench does about it. The stakes are small prizes and bragging rights, so most attackers are participants with a browser, a few friends and an evening, not nation states. Each entry says what stops the attack, where that lives in the code, and what doesn’t stop it.
People
- Participants want their project to win, want to see others’ work early, and want more time.
- Their friends will vote, and will make a few extra accounts if asked.
- Judges are mostly honest; a few favour a friend or want to see how their scores compare.
- Organizers are trusted with the event, but not with each other’s secrets or with rewriting history quietly.
- Anyone on the internet can reach the public pages and the API.
Attacks
Sybil voters: one person, many ballots
- Stops it: one ballot per account and per inbox, as unique constraints
on
portal_voter. Email addresses are normalized first (voting.normalize_email: lowercase,+tagdropped, Gmail dots andgooglemail.comfolded), soa.b+1@googlemail.comisab@gmail.com. A second spelling of an inbox gets no second ballot and is logged asvote.duplicate_refused; its link goes to the spelling just typed, and replaces the last one, so whoever typedalice+x@first doesn’t getalice@’s links. An address with a quoted local part ("a.b"@gmail.com) is refused rather than normalized. Accounts must confirm their address by a link before they can vote. Sign-ups, link requests and resets are rate limited per address (the written limits timesBALLOTBENCH_ADDRESS_LIMIT_SCALE, 10 by default, so a venue behind one NAT address isn’t locked out) and link requests per inbox (3 an hour). The organizer’s abuse panel (voting.abuse_report) flags networks (/24, or /64 for IPv6) with three or more voters, accounts created less than an hour before their first ballot, and three or more identical ballots. Organizers void a ballot with a reason; it leaves the tallies for good and the audit log says who did it and why. - Doesn’t stop it: someone with many real inboxes, a catch-all domain, or a botnet of addresses. There is no CAPTCHA, phone check or proof of personhood. Flags are never acted on automatically, because an office or a campus shares one network and friends vote alike; a person decides. Quadratic voting limits how much any one ballot can do, not how many there are.
Ballot stuffing: one ballot, more weight
- Stops it: the
vote_rulestrigger (migration 0012) locks the voter’s row and refuses any write that would take the sum of votes² pastvote_credits, so two tabs or a replayed request can’t overspend. It also refuses votes outside the window on the database clock, votes for a project that isn’t a submitted, non-duplicate project of the voter’s event (so no voting across events), votes for your own team, and votes by unconfirmed or voided voters. The app (voting.cast) checks the same things first so it can say why; it never trusts a cost sent by the client. Ballot writes are rate limited per address and per voter. - Doesn’t stop it: an email voter on a team whose members signed up with a different address. Own-team votes are refused for accounts in the database, and for email voters only when a team member’s address reaches the same inbox, which only the app can check.
Reading tallies or results early
- Stops it: tallies are shown only to the event’s organizers until
results are published (
voting.public_tallies),results.publishrefuses while voting is open, and published results are hidden again if voting reopens (results.visible_run), so nobody votes with the judges’ ranking in front of them. Unpublished results are a 404, not a 403. While voting is open organizers see how many ballots there are and what each voter spent, not the tallies, and a void can’t be undone: with live tallies and an undo, voiding one ballot, reading the tallies and counting it again would show what that voter chose. - Doesn’t stop it: an organizer telling people once voting closes. An organizer who voids a ballot after the close can see from the tallies what it held, at the price of throwing it away; one who moves the close date to read the tallies and then reopens voting can compare two reads. Both leave dated rows in the audit log.
Pushing a rival out as a duplicate
- Stops it: a flagged duplicate leaves judging and the ballot, so the
detector (
duplicates.find_duplicates) only compares submitted projects and calls the one submitted later the copy, bysubmitted_at, which only the server sets. A submit only ever flags the project being submitted (duplicates.flag_on_submit), so copying another team’s public title and repo into an old draft flags the copy, not the original. An organizer’s “it’s a different project” setsduplicate_cleared, and the detector never flags that project again. - Doesn’t stop it: a team that submits a placeholder early and edits it into a copy afterwards looks earlier to a whole-event scan. Submits never run one; only an organizer’s import does, and its flags are on the duplicates page for a person to judge.
Scraping drafts
- Stops it: every project query goes through
access.visible_projects: drafts are visible to their own team and the event’s organizers only, and anyone else gets a 404 on the page and the API. The gallery lists submitted projects only. - Doesn’t stop it: a team member sharing their screen, or a public repository linked from the draft.
Peeking at peer scores
- Stops it: judge queries start from
access.judge_assignments(user), so another judge’s review is a 404; naming another judge in/api/judge/scores?judge=is a 403, never an empty list. Review comments go to organizers only.scripts/isolation_curl.shtries all of this over HTTP. - Doesn’t stop it: judges talking to each other.
Judge collusion
- Stops it: conflict-of-interest triggers (
team_member_zz_not_judge,membership_judge_not_member,assignment_rules) keep a judge off their own team’s project and out of any team in an event they judge. Each project gets k reviews from different judges, and every review is in the audit log. - Doesn’t stop it: two judges who agree to push a project look like two judges who agree. Calibration can’t tell them apart (JUDGING.md).
Deadline gaming
- Stops it: the
project_deadlinetrigger (migration 0002) refuses creating, deleting or editing a project aftersubmissions_close, on the database’s clock, whatever path the write takes: page, API, admin or shell. The client’s clock is never asked. Moving the deadline is an auditedevent.update. - Doesn’t stop it: pushing to the linked repository after the deadline.
The portal records
submitted_at; checking the repository’s history against it is up to the judges.
Tampering with results or the audit log
- Stops it: published results are frozen to one calibration run, and
each run stores the SHA-256 of the exact scores it read (
input_digest). The audit log is append only (audit_readonly,audit_no_truncate) and hash chained;manage.py verify_auditnames the first row that doesn’t fit. Admin writes are audited like any other. - Doesn’t stop it: someone with the database superuser password can
disable the triggers and rewrite the whole chain from any point onwards.
Keep a copy of the latest
row_hashsomewhere else (the audit export has it) and a rewrite shows.
CSV formula injection
- Stops it: every exported cell that a spreadsheet would read as a
formula (starting with
=,+,-,@, tab or carriage return) and isn’t a plain number is prefixed with'(exports.cell), so a project called=HYPERLINK(...)opens as text. - Doesn’t stop it: a spreadsheet told to ignore that, or someone pasting cells into a formula by hand.
Comments as a weapon
- Stops it: bodies are plain text, escaped by Django’s autoescape (no
|safeanywhere), at most 2000 characters (checked by the serializer and by a database constraint), and rate limited to 10 per account per 10 minutes. The event’s organizers hide a comment and it disappears for everyone else; the row and the audit trail stay. - Doesn’t stop it: abuse that’s within the rules until an organizer reads it. There is no filter or pre-moderation.
Invite-link leakage
- Stops it: team and role invites are single use, only their hash is stored, team invites expire in 72 hours (or at the deadline) and role invites in 7 days, and both can be revoked. Voting links are single use, stored as a hash, and replaced by the next request. A voting link opened with GET only shows a button, so a mail scanner that fetches it doesn’t use it up.
- Doesn’t stop it: whoever gets a leaked link first. A leaked organizer
invite is a new organizer; check the Organizers list on the Settings page
and the
role.acceptrows in the audit log.
Token and session theft
- Stops it: API tokens are random, shown once, stored as SHA-256, and
revocable by their owner at
/me/tokens(or by an admin). Session cookies are HttpOnly and SameSite=Lax, every session POST needs a CSRF token, and no page can be framed except the read-only embed (below). Logins are limited per address (20 per person per 10 minutes, timesBALLOTBENCH_ADDRESS_LIMIT_SCALEfor a venue behind one NAT) and to 10 attempts per account per 10 minutes from anywhere, so a password can’t be guessed at network speed. - Doesn’t stop it: tokens don’t expire until revoked. Plain HTTP is the
default so the laptop demo works; behind TLS,
DJANGO_SECURE=1turns on secure cookies, HSTS and the redirect to HTTPS. A botnet still gets 10 guesses per account every 10 minutes; a strong password makes that useless.
Flooding
- Stops it: the rate limits above, counted in Postgres
(
ratelimit.allow) so every worker shares them. IPv6 is limited per /64, because one subscriber usually holds a whole /64. - Doesn’t stop it: a real denial of service. Put a proxy in front, and
set
BALLOTBENCH_TRUSTED_PROXIESto the number of proxies so the limits and the audit log see the visitor’s address (read from X-Forwarded-For, counting from the right, so a client can’t choose it), not the proxy’s.
Webhooks as a way into the private network
- Stops it: a webhook URL must be http(s) and every address its name
resolves to must be public (
webhooks.public:is_global, not multicast, not site-localfec0::/10; an IPv6 address carrying an IPv4 one, mapped, compatible, translated or NAT6464:ff9b::/96, is judged by the IPv4 one). That’s checked when the organizer saves it and again right before each send, and the connection goes to the address that was checked, so a DNS answer that changes between the check and the request (DNS rebinding) can’t redirect it. Redirects aren’t followed, and each delivery is signed with HMAC-SHA256 so the receiver can tell it’s from this portal. - Doesn’t stop it: an organizer choosing to deliver to their own network
with
BALLOTBENCH_WEBHOOKS_ALLOW_PRIVATE=1. Webhook payloads carry what the audit log carries for that event, so a webhook URL is a copy of the log: only organizers can add one, and adding one is itself audited. It keeps sending only while the organizer who added it (or who last resumed it) still organizes the event or is staff; otherwise the next change pauses it, with awebhook.pauseaudit row saying why.
A webhook receiver holding up the sender
- Stops it: one deadline of 10 seconds per attempt, for connecting,
sending and reading the answer together (a socket timeout alone counts
each read afresh, so a receiver sending a byte every few seconds could
hold the sender for ever). Only the status line is read, at most 16 KiB
looking for it, and never the body. Sends happen with no transaction or
row lock held: the sender claims a round of due deliveries, one per
webhook, by moving each
next_attempt_ata lease (2 minutes) ahead underSKIP LOCKED, commits, sends them all at once, then records each result. A sender that dies mid-send leaves its deliveries due again when the lease runs out. - Doesn’t stop it: a slow receiver still delays the other webhooks’ deliveries by up to one deadline per round, since a round waits for its slowest send. The DNS lookup before each send isn’t under the deadline; the system resolver’s own timeouts bound it.
The embed as a window in
- Stops it:
/embed/<slug>is the only page any site may frame (frame-ancestors *, no X-Frame-Options); everything else isDENY. It renders as an anonymous visitor whatever cookie or token comes with the request, so it never shows a draft or unpublished results, and it has no forms, so framing it can’t trick anyone into clicking something that changes state. - Doesn’t stop it: someone embedding a public gallery where you’d rather they didn’t. It’s public anyway.
Forged records and certificates
- Stops it: records are Ed25519 signatures over canonical JSON; the
public key is at
/.well-known/ballotbench-signing-key, and/verifychecks one without needing an account. Changing one character of a record makes it fail. Records never contain scores. The private key lives on the data volume, created with mode 0600, never in the repository. - Doesn’t stop it: someone who can read the data volume can sign anything. There’s no revocation list and no key rotation yet: a new key makes every old record fail to verify.
The API
The JSON API sits next to the pages and answers the same questions with the
same rules: every lookup goes through the same scoped querysets in
portal/access.py, and every write goes through the same database triggers.
Anything a page can tell you, the API can too, with one exception: the
organizer’s tools (settings, invitations, handing out reviews, calibration,
duplicates) are pages only.
Every example below is a real request against the demo stack on
http://localhost:8080, with the demo tokens from .dogfood.toml. The
read-only ones you can paste as they are. The ones that write change the
demo data, so run them on a stack you’re happy to reset with
docker compose down -v.
Authentication
Bearer tokens
Scripts authenticate with a token in the Authorization header:
curl -H "Authorization: Bearer bb_demo_judge_a_8d24f1" http://localhost:8080/api/judge/scores
A token is bb_ followed by 43 random characters. The portal stores only
its SHA-256, so a token is shown once, when it’s made, and a leaked database
doesn’t leak working tokens. A token acts as its user, with all of that
user’s roles; there are no scopes. A revoked token, or one whose user has
been deactivated, is refused.
The demo seed creates four tokens with fixed values, so the checker and this book can use them. They’re public, which is why a real deployment turns them off (see Running it for real).
| Who | Token | Roles |
|---|---|---|
| organizer | bb_demo_organizer_5c1e0a | organizer of both seeded events |
| judge_a | bb_demo_judge_a_8d24f1 | jdg_24 in the fixture event, judge of the open demo event |
| judge_b | bb_demo_judge_b_3a9e77 | jdg_29 in the fixture event |
| participant | bb_demo_participant_61b0c4 | team NorthKiln in the fixture event, participant in the demo event |
judge_a and judge_b share no project, so each has scores the other must
not see. Anyone signed in can issue and revoke their own tokens at
/me/tokens (linked from My events); a token is shown once and only its
hash is stored.
Browser sessions
A browser that’s signed in can call the API with its session cookie. Then
Django’s CSRF protection applies: a POST or PATCH needs the
X-CSRFToken header with the value of the csrftoken cookie, or it’s
refused with 403. Bearer requests need no CSRF token, because a browser
never attaches an Authorization header on its own, so a forged cross-site
request can’t carry one.
Pages accept tokens too
The HTML pages accept the same bearer tokens, for reading only. A GET
with a token sees exactly what that user would see in a browser, and no
session is created; this is what lets the isolation probe test the pages
with curl. Anything that would change something through a page (a POST)
is refused with 403 when it comes with a token, and so are /me/tokens and
the admin: a leaked token can’t mint more tokens. Scripts that write use the
API. A bad token on a page gets a plain-text 401.
Status codes
The same contract holds on the pages and in the API.
| Code | When | Example |
|---|---|---|
| 200, 201 | it worked; 201 when a project was created | |
| 400 | the request body is malformed: a missing required field, a wrong type, a track from another event | {"title": ["This field is required."]} |
| 401 | no credentials, or a bad token, on a route that needs them (API responses carry WWW-Authenticate: Bearer; pages send a browser to the login page instead) | {"detail": "invalid or revoked token"} |
| 403 | you’re signed in but your role can’t do this, or you named someone else’s data | {"detail": "judges can only read their own scores"} |
| 404 | the thing doesn’t exist, or it exists but isn’t yours to see: another judge’s assignment, another team’s draft, unpublished results | {"detail": "Not found."} |
| 409 | the window for this is closed, or a conflict of interest | {"detail": "Submissions for Sample Hack 2026 closed at 2026-03-01 18:00 UTC."} |
| 422 | the values are wrong: a score out of range, an unknown criterion, submitting a review with a criterion missing | {"detail": "Quality must be between 1 and 5"} |
The line between 403 and 404 is deliberate. If you may know a thing exists but can’t touch it (another team’s submitted project, which is in the public gallery), you get 403. If you may not even know it exists (another judge’s assignment, another team’s draft), you get 404, since a 403 would confirm it’s there.
A 409 or 422 can come from the app’s own check or from a database trigger.
The app checks first so it can give a clear message; the trigger is the
backstop, and access.guarded() turns its refusal into the same status
code. Errors are always JSON: {"detail": "..."}, or a field-by-field
object for a 400.
Endpoints
| Method | Path | Who may call it |
|---|---|---|
| GET | /api/events | anyone |
| GET | /api/events/<slug> | anyone |
| GET | /api/events/<slug>/projects | anyone; drafts only for their team and organizers |
| POST | /api/events/<slug>/projects | a member of a team in the event |
| GET | /api/projects/<id> | anyone who can see the project |
| PATCH | /api/projects/<id> | the project’s team |
| GET | /api/judge/assignments | judges |
| POST | /api/judge/assignments/<id>/review | the judge the assignment belongs to |
| GET | /api/judge/scores | judges, for themselves; organizers, for their events’ judges |
| GET | /api/events/<slug>/results | anyone once published; before that, the event’s organizers |
| GET | /api/events/<slug>/export/<kind>.csv | the event’s organizers |
| GET | /api/events/<slug>/results/signed | anyone once published: the ranking signed with the portal’s Ed25519 key |
| GET, POST | /api/events/<slug>/ballot | voters, in events that vote by signed-in account (confirmed address) |
| GET, POST | /api/projects/<id>/comments | anyone reads; signed-in users post |
| GET | /api/judge/record?event=<slug> | a judge, for their own signed record; organizers may name a judge of their event |
| GET | /api/participant/record?event=<slug> | a participant, for their own signed record |
| POST | /api/verify | anyone: checks a signed record or signed results |
| GET | /.well-known/ballotbench-signing-key | anyone: the public key records and results are signed with |
| GET | /api/events/<slug>/export/bundle.json | the event’s organizers: the whole event as one file |
| POST | /api/events/import?slug=<new> | site admins: a bundle back in as a new event |
| GET | /api/schema | anyone |
| GET | /api/docs | anyone |
Site admins (is_staff) count as organizers of every event.
Events
List events
GET /api/events: every event, the latest deadline first, with its
windows, tracks and rubric. No authentication.
curl http://localhost:8080/api/events
[
{
"slug": "demo-open",
"name": "Demo Hack (open)",
"...": "..."
},
{
"slug": "sample-hack-2026",
"name": "Sample Hack 2026",
"description": "",
"submissions_open": "2026-02-26T00:00:00Z",
"submissions_close": "2026-03-01T18:00:00Z",
"judging_open": "2026-03-01T18:00:00Z",
"judging_close": null,
"voting_open": null,
"voting_close": null,
"results_published_at": null,
"reviews_per_project": 3,
"max_team_size": 4,
"tracks": [{"id": 1, "name": "Developer tools"}, {"id": 2, "name": "Data and analytics"}, "..."],
"criteria": [
{"key": "functionality", "name": "Functionality", "weight": "1.000", "min_value": 1, "max_value": 5},
{"key": "quality", "name": "Quality", "weight": "1.000", "min_value": 1, "max_value": 5},
{"key": "innovation", "name": "Innovation", "weight": "1.000", "min_value": 1, "max_value": 5}
]
}
]
All times are UTC. A null close means the window never closes. Weights are
decimals, sent as strings so they don’t lose precision.
One event
GET /api/events/<slug>: the same object for one event, or 404.
curl http://localhost:8080/api/events/sample-hack-2026
Projects
A project looks like this in every response:
{
"id": 1,
"event": "sample-hack-2026",
"team": "NorthKiln",
"track": 4,
"title": "Glass Signal",
"tagline": "",
"summary": "One line of what it does.",
"description": "",
"repo_url": "https://example.org/repo/01",
"demo_url": "",
"tags": [],
"status": "submitted",
"submitted_at": "2026-02-27T04:08:00Z",
"duplicate_of": null
}
track is a track id from the event. status is draft or submitted.
duplicate_of is the id of the earlier project this one repeats, if it’s
been flagged; flagged projects stay listed but aren’t ranked.
List an event’s projects
GET /api/events/<slug>/projects: the event’s submitted projects, in id
order, for anyone. Signed in, you also see your own team’s drafts, and an
organizer sees every draft in the event.
curl http://localhost:8080/api/events/sample-hack-2026/projects
On the fixture that’s 41 projects, including the duplicate prj_41 (id 41,
"duplicate_of": 7).
One project
GET /api/projects/<id>: one project, if you can see it. Another team’s
draft is a 404, not a 403.
curl http://localhost:8080/api/projects/1
Create a project
POST /api/events/<slug>/projects: creates a project for your team in that
event. You need to be on a team in the event; teams are made on the event’s
page, since the API has no team routes.
| Field | Type | Notes |
|---|---|---|
title | string | required |
tagline, summary, description | string | optional |
repo_url, demo_url | URL | optional |
tags | list of strings | trimmed, lowercased and de-duplicated; at most 10 |
track | track id | must be a track of this event |
submit | boolean | true submits it; otherwise it’s saved as a draft |
| Outcome | Code |
|---|---|
| created | 201, with the project |
| not signed in | 401 |
| no team in this event | 403 you need a team in this event to submit a project |
| before the window opens or after it closes | 409, with the time |
| a bad field | 400 |
Each call makes a new project; a team can hold more than one (the fixture’s team CopperLedger has two). When a project is submitted, the portal checks it against the event’s earlier projects and flags it if it repeats one.
This is the request the acceptance checker makes. The fixture event closed on 1 March 2026, so it’s refused:
curl -X POST http://localhost:8080/api/events/sample-hack-2026/projects \
-H "Authorization: Bearer bb_demo_participant_61b0c4" \
-H "Content-Type: application/json" \
-d '{"title": "late", "summary": "x"}'
{"detail": "Submissions for Sample Hack 2026 closed at 2026-03-01 18:00 UTC."}
with status 409. On the open demo event, once the participant has started a team on the event’s page, the same call creates a draft:
curl -X POST http://localhost:8080/api/events/demo-open/projects \
-H "Authorization: Bearer bb_demo_participant_61b0c4" \
-H "Content-Type: application/json" \
-d '{"title": "Night Owl Radio", "tagline": "Offline-first radio for field teams",
"summary": "Mesh radio notes.", "repo_url": "https://example.org/night-owl",
"tags": ["radio", "mesh"]}'
{
"id": 42,
"event": "demo-open",
"team": "Night Owls",
"track": null,
"title": "Night Owl Radio",
"tagline": "Offline-first radio for field teams",
"summary": "Mesh radio notes.",
"description": "",
"repo_url": "https://example.org/night-owl",
"demo_url": "",
"tags": ["radio", "mesh"],
"status": "draft",
"submitted_at": null,
"duplicate_of": null
}
Edit or submit a project
PATCH /api/projects/<id>: changes the fields you send and leaves the rest.
Only members of the project’s team may; send "submit": true to submit it.
submitted_at is set from the database’s clock, and there’s no way to
unsubmit through the API.
| Outcome | Code |
|---|---|
| saved | 200, with the project |
| you can’t see the project | 404 |
| not signed in | 401 |
| you can see it but it isn’t your team’s | 403 only the project's team can edit it |
| the submission window is closed | 409 |
| a bad field | 400 |
curl -X PATCH http://localhost:8080/api/projects/42 \
-H "Authorization: Bearer bb_demo_participant_61b0c4" \
-H "Content-Type: application/json" \
-d '{"demo_url": "https://example.org/night-owl/demo", "submit": true}'
{"id": 42, "...": "...", "demo_url": "https://example.org/night-owl/demo",
"status": "submitted", "submitted_at": "2026-09-28T05:13:51.891600Z", "duplicate_of": null}
After the deadline, editing your own project is a 409, and editing another
team’s is a 403 (prj_02 is another team’s):
curl -X PATCH http://localhost:8080/api/projects/1 \
-H "Authorization: Bearer bb_demo_participant_61b0c4" \
-H "Content-Type: application/json" -d '{"title": "late"}'
# 409 {"detail": "Submissions for Sample Hack 2026 closed at 2026-03-01 18:00 UTC."}
curl -X PATCH http://localhost:8080/api/projects/2 \
-H "Authorization: Bearer bb_demo_participant_61b0c4" \
-H "Content-Type: application/json" -d '{"title": "mine"}'
# 403 {"detail": "only the project's team can edit it"}
Judging
Your assignments
GET /api/judge/assignments: your own assignments in every event you
judge, and nobody else’s. 403 judges only if you judge no event.
curl -H "Authorization: Bearer bb_demo_judge_a_8d24f1" http://localhost:8080/api/judge/assignments
[
{"id": 16, "event": "sample-hack-2026", "project": 6, "project_title": "Dry Compass", "status": "done"},
{"id": 38, "event": "sample-hack-2026", "project": 12, "project_title": "Open Beacon", "status": "done"},
"..."
]
status is pending, done or recused. Stepping aside from a project
(recusal) is on the scoresheet page, not in the API.
Score an assignment
POST /api/judge/assignments/<id>/review: saves your scores for one of your
own assignments.
| Field | Type | Notes |
|---|---|---|
scores | object | criterion key to an integer in that criterion’s range |
comment | string | optional, up to 5000 characters; organizers see it, other judges never do |
submit | boolean | true submits; otherwise it’s a draft. Submitting needs every criterion. |
| Outcome | Code |
|---|---|
| saved | 200 |
| not your assignment, whatever the id | 404 |
| judging hasn’t opened, or has closed | 409, with the time |
| you stepped aside from this project | 409 |
| a score out of range, an unknown criterion, or submitting with one missing | 422 |
| a score that isn’t an integer | 400 |
You can save and resubmit as often as you like while judging is open. Every save is in the audit log with the scores before and after.
curl -X POST http://localhost:8080/api/judge/assignments/16/review \
-H "Authorization: Bearer bb_demo_judge_a_8d24f1" \
-H "Content-Type: application/json" \
-d '{"scores": {"functionality": 3, "quality": 4, "innovation": 5}, "comment": "Clear demo.", "submit": true}'
{"assignment": 16, "submitted_at": "2026-09-28T05:13:51.974376Z", "weighted_total": 0.75}
weighted_total is the review’s score on the 0 to 1 scale (see
the method). The fixture event’s
judging window has no close, so this really does change the fixture’s
scores, and with them the calibration fingerprint.
The same assignment, as another judge, doesn’t exist:
curl -X POST http://localhost:8080/api/judge/assignments/16/review \
-H "Authorization: Bearer bb_demo_judge_b_3a9e77" \
-H "Content-Type: application/json" -d '{"scores": {"quality": 5}}'
# 404 {"detail": "No JudgeAssignment matches the given query."}
Your scores
GET /api/judge/scores: your own reviews, drafts included, with every
criterion’s score and the weighted total.
curl -H "Authorization: Bearer bb_demo_judge_a_8d24f1" http://localhost:8080/api/judge/scores
[
{
"assignment": 16,
"event": "sample-hack-2026",
"project": 6,
"project_title": "Dry Compass",
"judge": "diego.herrera@example.org",
"status": "done",
"submitted_at": "2026-03-01T18:00:00Z",
"comment": "Runs clean.",
"scores": {"functionality": 2, "quality": 3, "innovation": 5},
"weighted_total": 0.5833333333333334
},
"..."
]
Two query parameters:
event=<slug>keeps one event’s reviews.judge=<judge>names a judge, by fixture id (jdg_24), email or user id. A judge may name only themselves. Naming anyone else is a 403, never an empty list, so a refusal can’t be mistaken for “no scores”; and a name that matches nobody is a 403 too, so the answer doesn’t reveal which judges exist. An organizer may name any judge and gets that judge’s reviews in the events they organize (404no such judgeif the name matches nobody).
| Caller | Result |
|---|---|
a judge, no judge= | 200, their own reviews |
| a judge naming another judge, or nobody | 403 judges can only read their own scores |
someone who judges nothing, no judge= | 403 judges only |
| an organizer naming a judge | 200, that judge’s reviews in the organizer’s events |
| no credentials | 401 |
The acceptance checker’s peer probe:
curl -H "Authorization: Bearer bb_demo_judge_b_3a9e77" "http://localhost:8080/api/judge/scores?judge=jdg_24"
# 403 {"detail": "judges can only read their own scores"}
curl -H "Authorization: Bearer bb_demo_organizer_5c1e0a" \
"http://localhost:8080/api/judge/scores?judge=jdg_24&event=sample-hack-2026"
# 200, jdg_24's eleven fixture reviews
Results
GET /api/events/<slug>/results: the ranked results. Until they’re
published this is a 404 for everyone except the event’s organizers, who get
the latest calibration run (with "published_at": null). After publishing,
everyone gets the published run, and later runs change nothing here.
curl http://localhost:8080/api/events/sample-hack-2026/results
# 404 {"detail": "Not found."} until an organizer publishes
curl -H "Authorization: Bearer bb_demo_organizer_5c1e0a" http://localhost:8080/api/events/sample-hack-2026/results
{
"event": "sample-hack-2026",
"published_at": null,
"run": 1,
"input_digest": "397dd35f376eb052b0dfcd8fda8309f46bad4771e9caa4e78035175807178f4a",
"method": "offset-scale-noise/v1",
"projects": [
{
"rank": 1,
"raw_rank": 31,
"project": 7,
"title": "Dry Harbour",
"team": "CopperLedger",
"calibrated": 0.8321,
"se": 0.0304,
"raw_mean": 0.5833,
"reviews": 5,
"rank_interval": [1, 39]
},
"..."
]
}
That’s the fixture after one calibration run and before publishing. Only
ranked projects are listed; duplicates and projects with no usable reviews
are left out. rank_interval is the 90% bootstrap range the page shows as
“could be”; se is the model’s own standard error, which is narrower
because it treats every judge’s habits as known. input_digest is the
fingerprint of the scores the run read (how to check
it).
CSV exports
GET /api/events/<slug>/export/<kind>.csv: one export per stage of the
event, for its organizers. 401 without credentials, 403 for anyone else,
404 for a kind that doesn’t exist. Every download writes an export.csv
row to the audit log.
curl -H "Authorization: Bearer bb_demo_organizer_5c1e0a" \
http://localhost:8080/api/events/sample-hack-2026/export/scores.csv | head -3
review_id,judge_email,project_id,project,submitted_at,functionality,quality,innovation,weighted_0_1,comment
1,marek.nowak@example.org,1,Glass Signal,2026-03-01T18:00:00+00:00,2,4,2,0.4167,Runs clean.
3,pavel.ivanov@example.org,1,Glass Signal,2026-03-01T18:00:00+00:00,2,5,3,0.5833,Solid.
| Kind | One row per | Columns |
|---|---|---|
registrations | membership | user_email, name, role, tracks |
teams | team member | team_id, external_id, team, member_email, joined_at |
projects | project, drafts included | project_id, external_id, title, team, track, status, submitted_at, repo_url, demo_url, tags, duplicate_of |
assignments | assignment | assignment_id, judge_email, project_id, project, status, source, created_at |
scores | review, drafts included | review_id, judge_email, project_id, project, submitted_at, one column per criterion, weighted_0_1, comment |
results | project in the latest run | run, rank, rank_low, rank_high, raw_rank, project_id, project, reviews, raw_mean, calibrated, se, excluded, input_digest |
judges | judge in the latest run | run, judge_email, reviews, offset, scale, noise, flag |
audit | audit row for the event | seq, ts, actor, action, object_type, object_id, ip, before, after, prev_hash, row_hash |
The file is served as text/csv with a download name like
sample-hack-2026-scores.csv. Any cell that starts with =, +, - or
@ (and isn’t a number) gets a leading ', so a spreadsheet won’t run a
project title as a formula. results and judges always describe the
latest run, published or not; the run column says which.
The OpenAPI document and the reference page
GET /api/schema serves the OpenAPI 3.0 document, generated by
drf-spectacular from the same views, so it can’t drift from the code. It’s
YAML by default; ask for JSON with ?format=json or Accept: application/json.
curl http://localhost:8080/api/schema # YAML
curl "http://localhost:8080/api/schema?format=json" # JSON
GET /api/docs is a reference page rendered on the server from that same
document. It needs no JavaScript and no CDN, so it works on a stack with no
network. Both are public.
The schema lists each route’s success response only. The error codes are in this chapter and in each route’s description.
Running it for real
The image that runs the demo is the image you run an event on. What changes is configuration: a real secret, your hostname, no demo accounts, TLS in front, and backups. This chapter covers each, then the maintenance commands and the few sharp edges I know about.
A checklist
For an event that people outside your laptop will use:
- Start from an empty database, with the demo accounts off (below).
- Put your settings in a
docker-compose.override.yml(below) rather than editingdocker-compose.yml. - Create the first site admin with
createsuperuser. - Put a reverse proxy with TLS and rate limits in front, and set
DJANGO_SECURE=1. - Schedule
pg_dumpandverify_audit. - Before the event, run through the tour once on your own deployment.
Environment variables
Settings come from the environment, read in
src/ballotbench/settings.py,
src/entrypoint.sh
and the seed command. This is all of them.
| Variable | Default | In docker-compose.yml | What it does |
|---|---|---|---|
DJANGO_SECRET_KEY | none | not set | signs sessions and CSRF tokens; wins over the file below |
DJANGO_SECRET_KEY_FILE | none | /data/secret_key | a file holding the key; the entrypoint writes a random one there on first boot if it’s missing or empty |
DJANGO_DEBUG | off | not set | 1 shows Django’s debug pages and allows a built-in insecure key; for development only |
DJANGO_ALLOWED_HOSTS | localhost,127.0.0.1,[::1] | adds web | comma-separated host names the portal answers to |
DJANGO_SECURE | off | not set | 1 behind a TLS proxy: secure cookies, HTTPS redirect, HSTS, trust the proxy’s scheme header (below) |
POSTGRES_DB | ballotbench | not set | database name |
POSTGRES_USER | ballotbench | not set | database user |
POSTGRES_PASSWORD | ballotbench | ballotbench | database password |
POSTGRES_HOST | localhost | db | database host |
POSTGRES_PORT | 5432 | not set | database port |
BALLOTBENCH_DEMO_SEED | off | "1" | 1 loads the fixture event, the open demo event and the demo accounts with their fixed tokens on boot; anything else loads nothing |
BALLOTBENCH_FIXTURES | /app/fixtures.json in the image | not set | the fixture file the seed imports on every boot |
BALLOTBENCH_TRUSTED_PROXIES | 0 | not set | how many reverse proxies sit in front; the caller’s address is then read from X-Forwarded-For, counting from the right |
BALLOTBENCH_ADDRESS_LIMIT_SCALE | 10 | not set | multiplies every per-address rate limit, so a venue behind one NAT address isn’t locked out; per-account limits aren’t scaled |
DJANGO_EMAIL_BACKEND | the database outbox | not set | Django’s SMTP backend (django.core.mail.backends.smtp.EmailBackend, plus the usual EMAIL_* settings) to really send mail |
WEB_WORKERS | 3 | not set | gunicorn worker processes |
If neither DJANGO_SECRET_KEY nor DJANGO_SECRET_KEY_FILE is set, and debug
is off, the portal refuses to start rather than run with a guessable key.
The generated key lives on the webdata volume, so sessions survive a
restart and the key never lives in the repository.
The db service has its own POSTGRES_DB, POSTGRES_USER and
POSTGRES_PASSWORD, read by the Postgres image. The web service’s values
must match them. The Postgres image only reads them when it creates the
database, the first time the pgdata volume is used; changing the password
later means ALTER USER inside the database as well.
DJANGO_ALLOWED_HOSTS must keep 127.0.0.1: the container’s health check
asks for http://127.0.0.1:8080/projects, and Django refuses a host it
doesn’t know.
An override file
Compose reads docker-compose.override.yml next to docker-compose.yml
automatically, so your settings stay out of the file you pull updates into.
For a deployment at judging.example.org behind a proxy on the same
machine:
# docker-compose.override.yml
services:
db:
environment:
POSTGRES_PASSWORD: a-long-random-password
web:
environment:
POSTGRES_PASSWORD: a-long-random-password
DJANGO_ALLOWED_HOSTS: judging.example.org,localhost,127.0.0.1
DJANGO_SECURE: "1"
BALLOTBENCH_DEMO_SEED: "0"
# Only the proxy on this machine talks to the portal.
ports: !override
- "127.0.0.1:8080:8080"
# With DJANGO_SECURE=1 a plain-http request is redirected to https,
# so the health check has to say it came through the proxy.
healthcheck:
test: ["CMD", "python", "-c", "import urllib.request as u; u.urlopen(u.Request('http://127.0.0.1:8080/projects', headers={'X-Forwarded-Proto': 'https'}), timeout=3)"]
!override replaces the port list instead of adding to it; it needs Docker
Compose 2.24 or newer. Check what Compose will actually run with
docker compose config.
Demo accounts
With BALLOTBENCH_DEMO_SEED=1, every boot makes sure these exist: the site
admin admin@ballotbench.local, the organizer organizer@ballotbench.local,
two fixture judges and a fixture participant, all with the password
ballotbench-demo, and four API tokens whose values are printed in the
README. Everything about them is public, so a real deployment must not have
them.
- On a new deployment, set
BALLOTBENCH_DEMO_SEEDto0before the first boot. None of them is created, including the admin, so create your own withcreatesuperuser. - If they already exist, setting the variable to
0stops the seed creating them; it doesn’t delete them. Start again from an empty volume (docker compose down -v, which deletes everything), or switch them off from a shell:
docker compose exec web python manage.py shell -c "
from django.utils import timezone
from portal.models import ApiToken, User
ApiToken.objects.filter(label='demo', revoked_at=None).update(revoked_at=timezone.now())
User.objects.filter(email__in=['admin@ballotbench.local', 'organizer@ballotbench.local',
'diego.herrera@example.org', 'ines.rocha@example.org', 'priya1@example.org']).update(is_active=False)
"
A deactivated user can’t sign in, and their tokens are refused.
The seeded events follow the same switch. With BALLOTBENCH_DEMO_SEED set
to anything but 1, the seed loads nothing: no fixture event, no demo event,
no accounts. A deployment that already has them keeps them (the seed never
deletes); remove them for good with
docker compose exec web python manage.py delete_event sample-hack-2026 --yes
(and demo-open), and they won’t come back.
TLS and a reverse proxy
The portal speaks plain HTTP on port 8080. For anything public, put a
reverse proxy in front that terminates TLS, and set DJANGO_SECURE=1. That
turns on:
SESSION_COOKIE_SECUREandCSRF_COOKIE_SECURE: cookies only over HTTPS;SECURE_PROXY_SSL_HEADER = ("HTTP_X_FORWARDED_PROTO", "https"): a request counts as HTTPS when the proxy says so inX-Forwarded-Proto;SECURE_SSL_REDIRECT: plain-HTTP requests are redirected to HTTPS;SECURE_HSTS_SECONDSof 30 days: browsers stay on HTTPS;CSRF_TRUSTED_ORIGINS:https://plus each host inDJANGO_ALLOWED_HOSTS.
It’s off by default so the demo works on http://localhost. With it on,
the acceptance checker and the isolation probe, which talk plain HTTP to
port 8080, get redirects instead of answers; they’re for the demo stack.
Trusting X-Forwarded-Proto is only safe if the proxy always sets it
itself and nobody can reach port 8080 except through the proxy. Bind the
port to 127.0.0.1 as in the override above, or keep it off the host
entirely.
The portal limits logins per address and per account, and sign-ups and
password resets per address (in the database, so every worker shares the
counts; see BALLOTBENCH_ADDRESS_LIMIT_SCALE). A proxy is still the right
place for a coarse limit on everything, before a request reaches Python. An
nginx example:
limit_req_zone $binary_remote_addr zone=bb_auth:10m rate=10r/m;
limit_req_zone $binary_remote_addr zone=bb_api:10m rate=10r/s;
server {
listen 80;
server_name judging.example.org;
return 301 https://$host$request_uri;
}
server {
listen 443 ssl;
server_name judging.example.org;
ssl_certificate /etc/letsencrypt/live/judging.example.org/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/judging.example.org/privkey.pem;
client_max_body_size 1m;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
location ~ ^/(login|signup)$ {
limit_req zone=bb_auth burst=5 nodelay;
proxy_pass http://127.0.0.1:8080;
}
location /api/ {
limit_req zone=bb_api burst=20;
proxy_pass http://127.0.0.1:8080;
}
location / {
proxy_pass http://127.0.0.1:8080;
}
}
The limits are a starting point: ten sign-in or sign-up attempts a minute per address is plenty for a person and slow for a script. A whole venue behind one NAT address shares that budget, so watch the proxy’s log on the day.
The audit log records the address a request came from, as the portal sees
it. It deliberately doesn’t trust X-Forwarded-For, so behind a proxy the
address column shows the proxy’s address, not the visitor’s; the proxy’s
own access log has the real one.
The portal sends no email. Invitation links are shown once, to whoever
makes them, to pass on however they like; that’s what lets it run with no
network. EMAIL_BACKEND is Django’s console backend, so anything Django
itself tries to send is written to the web container’s log.
If you add email, it’s a change in settings.py, since there are no email
variables today. Replace the EMAIL_BACKEND line with something like:
EMAIL_BACKEND = "django.core.mail.backends.smtp.EmailBackend"
EMAIL_HOST = env("EMAIL_HOST", "localhost")
EMAIL_PORT = int(env("EMAIL_PORT", "587"))
EMAIL_HOST_USER = env("EMAIL_HOST_USER", "")
EMAIL_HOST_PASSWORD = env("EMAIL_HOST_PASSWORD", "")
EMAIL_USE_TLS = True
DEFAULT_FROM_EMAIL = env("DEFAULT_FROM_EMAIL", "ballotbench@judging.example.org")
and add those variables to your override file. Keep in mind that the stack then needs the network to reach the mail server.
Backups
All state is in two Docker volumes: pgdata (the database) and webdata
(the generated secret key). The database is plain Postgres, so pg_dump
works. Take a custom-format dump from the running stack:
docker compose exec -T db pg_dump -U ballotbench -Fc ballotbench > ballotbench-$(date +%Y%m%d-%H%M).dump
-T matters: without it Compose allocates a terminal and can mangle the
binary output. I’d run this from cron every hour during an event, and keep
the files off the machine.
To restore, stop the portal, recreate the database empty, load the dump, start the portal and check the audit chain:
docker compose stop web
docker compose exec -T db dropdb -U ballotbench ballotbench
docker compose exec -T db createdb -U ballotbench ballotbench
docker compose exec -T db pg_restore -U ballotbench -d ballotbench --no-owner < ballotbench-20260929-1800.dump
docker compose start web
docker compose exec web python manage.py verify_audit
Restore into an empty database, not over the live one. pg_restore loads
the rows before it creates the triggers, so the audit chain, the deadline
and the other rules come back exactly as they were, and verify_audit
should report the chain intact. Losing webdata only signs everyone out;
the entrypoint makes a new key.
Upgrades
docker compose exec -T db pg_dump -U ballotbench -Fc ballotbench > before-upgrade.dump
git pull
docker compose up -d --build
docker compose logs -f web
There’s no separate migration step. Every boot runs
manage.py migrate --noinput, then the seed, then gunicorn, so a new
version’s migrations (including new triggers) are applied when the new
container starts; with no new migrations it’s a no-op. If an upgrade goes
wrong, restore the dump you just took rather than trying to migrate
backwards.
The Postgres image is pinned (postgres:18.4). A patch release is a change
of tag; a new major version needs a dump and a restore into a fresh volume,
as with any Postgres.
Management commands
Run them with docker compose exec web python manage.py <command>.
seed
Imports the fixture event and creates the open demo event, and with
BALLOTBENCH_DEMO_SEED=1 the demo accounts. The entrypoint runs it on every
boot. It’s idempotent: imported rows are keyed by their fixture ids, so a
second run creates nothing and never overwrites what people changed. You
shouldn’t need to run it by hand.
To load another event in the fixture’s shape, use the importer from a shell and give it its own slug:
docker compose cp other-event.json web:/tmp/other-event.json
docker compose exec web python manage.py shell -c "
import json
from portal.importer import import_event
event, counts = import_event(json.load(open('/tmp/other-event.json')), slug='other-event')
print(event.slug, counts)
"
seed --fixtures other-event.json looks like the way to do this, but it
names every imported event sample-hack-2026, so it fails once the fixture
event exists.
createsuperuser
Creates a site admin, who can do everything an organizer can in every event
and use the Django admin at /admin/. You need one when the demo accounts
are off.
docker compose exec web python manage.py createsuperuser
It asks for an email address and a password. From a script:
docker compose exec -e DJANGO_SUPERUSER_PASSWORD='a long passphrase' web \
python manage.py createsuperuser --noinput --email you@example.org
verify_audit
Recomputes the audit log’s hash chain from the first row to the last and
names the first row that doesn’t fit: a missing row, a changed prev_hash,
or a row edited after it was written. It exits with status 1 if the chain is
broken, so it can run from cron or CI.
audit chain intact: 26 rows, head e1cd7b2acec59727
Run it after a restore, before publishing results, and on a schedule. The database refuses edits to the log from the app and from ordinary SQL; this catches the one thing that can get past that, a database superuser.
normalization_proof
Recomputes every calibration claim in the method chapter from an event’s scores: which judges carry no weight and why, the agreement test, the ranking with its intervals, the invariance checks on these scores, and a benchmark on synthetic events. It takes a few seconds and is deterministic.
docker compose exec web python manage.py normalization_proof # the fixture event
docker compose exec web python manage.py normalization_proof --event your-event # yours
Use it when you want to see the model at work on your own event before you publish, or to check the book’s numbers. It needs an event with reviews.
delete_event
Deletes an event and everything in it. After submissions close, the
database refuses to delete submitted projects, which is what you want
during an event and in the way afterwards; this command is the one
deliberate way round it. It asks for --yes, and it writes an
event.delete row to the audit log first. The audit rows of the deleted
event stay, since they aren’t tied to it by a foreign key.
docker compose exec web python manage.py delete_event old-hack-2025 --yes
The maintenance flag
Some database rules would stop legitimate maintenance. The fixture’s projects were submitted before a close date that has already passed, so importing them is, to the deadline trigger, a late submission. Deleting a closed event means deleting projects after the deadline. The fixture’s team sizes aren’t checked against the event’s limit either.
For those cases the triggers look at a session setting,
ballotbench.import. When it’s on, the deadline trigger and the
team-size check step aside. Only two pieces of code set it, the fixture
importer and delete_event, and both use SET LOCAL, so it lasts only
until their own transaction ends, and both write an audit row saying what
they did.
It’s not a security boundary. Anyone with SQL access to the database can set
it, or drop a trigger outright. The triggers are there to stop the app, the
admin and a careless shell from breaking the rules by accident. Deliberate
changes by someone with database access are what the audit chain and
verify_audit are for. The app connects as the user the Postgres image
creates, which is a superuser, so keep database access to the people who
run the event.
API tokens
Anyone signed in issues and revokes their own tokens at /me/tokens; a
token acts with that person’s roles and nothing more, is shown once, and
only its hash is stored. An admin can also issue one for any account from a
shell:
docker compose exec web python manage.py shell -c "
from portal.auth import issue_token
from portal.models import User
print(issue_token(User.objects.get(email='organizer@example.org'), 'results script'))
"
bb_YC5-kg1ErFIezXIBwN0h1Y6ysq9uxVNkEDkaB2_IEmM
The label is for you, to tell tokens apart. The token has all of its user’s roles. To revoke it, set its revoked time, from the Django admin (Api tokens) or from a shell:
docker compose exec web python manage.py shell -c "
from django.utils import timezone
from portal.models import ApiToken
ApiToken.objects.filter(user__email='organizer@example.org', label='results script',
revoked_at=None).update(revoked_at=timezone.now())
"
A token issued or revoked from a shell isn’t in the audit log; one revoked through the admin is.
With the network off
Once the images are built, the stack needs no network. The override
docker-compose.offline.yml puts both containers on a Docker network with
no route out:
docker compose down -v
docker compose -f docker-compose.yml -f docker-compose.offline.yml up
Docker doesn’t publish ports from an internal network, so in this mode the portal isn’t reachable from the host’s browser. Check it from inside, as CI does:
docker compose exec web python -c "import urllib.request as u; [u.urlopen('http://localhost:8080' + p) for p in ('/projects', '/static/portal/site.css')]"
docker compose up --force-recreate puts the containers back on the normal
network.
Known sharp edges
Things I’d want to know before running an event on it:
- An account must confirm its address before it can vote, but anyone with a working inbox can sign up. Logins, sign-ups and resets are rate limited; a proxy with its own limits is still wise.
- Behind a proxy, set
BALLOTBENCH_TRUSTED_PROXIESto the number of proxies, or the audit log and the rate limits see the proxy’s address. - Mail goes to the outbox table, readable in the admin. For real delivery,
set
DJANGO_EMAIL_BACKENDto Django’s SMTP backend and configure it. - Calibration runs inside the organizer’s request. On the fixture that’s a few seconds; calibration needs no background worker.
- Community vote tallies are computed when read. Voiding a ballot after publishing changes the public numbers (and is in the audit log).
The acceptance checker
Dogfood’s organizers give every team the same checker,
run.py: a
standard-library Python script that reads a team’s .dogfood.toml, makes
seven HTTP requests against the running portal, and prints a report of
which tiers it could verify. This chapter explains the contract as
ballotbench meets it, and the stricter probe I wrote to go past it.
.dogfood.toml, line by line
[portal]
base_url = "http://localhost:8080"
[tiers]
claimed = ["T1", "T2"]
pitch = "A self-hosted hackathon portal whose judging you can defend: weighted rubrics, calibrated judges, and deadlines and isolation enforced in the database."
[auth]
organizer = "Authorization: Bearer bb_demo_organizer_5c1e0a"
judge_a = "Authorization: Bearer bb_demo_judge_a_8d24f1"
judge_b = "Authorization: Bearer bb_demo_judge_b_3a9e77"
participant = "Authorization: Bearer bb_demo_participant_61b0c4"
[routes]
gallery = "/projects"
submit = "/api/events/sample-hack-2026/projects"
judge_scores = "/api/judge/scores"
peer_scores = "/api/judge/scores?judge=jdg_24"
csv_export = "/api/events/sample-hack-2026/export/scores.csv"
| Line | Meaning |
|---|---|
base_url | where the checker sends every request: the port docker compose up publishes |
claimed | the tiers this entry claims: T1 (submissions and the gallery) and T2 (judging and isolation), the ones run.py has checks for. The public vote (T3) and the stretch pieces (T4) are built and tested, but the checker can’t verify them, so they aren’t claimed. |
pitch | a one-line description of the entry; run.py doesn’t read it |
organizer, judge_a, judge_b, participant | a complete HTTP header for each role. The checker splits it at the first : and sends it as is. These are the demo tokens the seed creates when BALLOTBENCH_DEMO_SEED=1. |
gallery | the public page listing submitted projects |
submit | where a participant creates a project, here in the fixture event, which is closed |
judge_scores | where a judge reads their own scores |
peer_scores | the URL that would return judge A’s scores: jdg_24 is judge A’s id in the fixture |
csv_export | an organizer’s CSV export |
The four accounts are chosen so the checks mean something: judge_a
(jdg_24) and judge_b (jdg_29) both have fixture reviews and share no
project, so each has scores the other must not see, and the participant is
on a real fixture team (NorthKiln), so the late submission is refused for
being late, not for having no team.
The seven checks
| Tier | Check | The request | Passes when | What answers it |
|---|---|---|---|---|
| T1 | gallery is public | GET /projects, no header | 200 | the gallery view, which lists submitted projects to anyone |
| T1 | project from fixtures shown | the same response | it contains the title of one of the fixture’s first three projects | the seed imports the fixture on boot, and the gallery orders page one by id, so Glass Signal, Small Meadow and Deep Compass lead |
| T1 | closed event refuses submissions | POST to submit as participant, with a title and summary | any 4xx | 409 with the close time: the API checks the window on the database’s clock, and the project_deadline trigger would refuse it anyway |
| T2 | judge sees own scores | GET judge_scores as judge_a | 200 | /api/judge/scores returns the caller’s own reviews |
| T2 | judge cannot see peer scores | GET peer_scores as judge_b | 401 or 403 | 403: a judge may only name themselves, and gets a refusal, never an empty list |
| T2 | participant blocked | GET judge_scores as participant | 401 or 403 | 403 judges only |
| T2 | csv export works | GET csv_export as organizer | 200, and the first line has a comma | the scores export, whose first line is the CSV header |
The checker is lenient in places (any 4xx for the late submission, 401 or 403 for the peer probe). ballotbench answers with the specific code in each case; the API chapter has the contract.
A tier counts as verified only if every one of its checks passes and every
tier below it is verified too. run.py has checks for T1 and T2 only, so
those are the most it can verify.
Regenerating acceptance-report.txt
The report in the repository is the checker’s output against the Docker build on a fresh volume. To make it again, from the repository root:
docker compose down -v
docker compose up -d --build --wait
python3 run.py .dogfood.toml > acceptance-report.txt
Run it from the root so it finds fixtures.json (it looks in the current
directory, next to run.py, next to the config, and in a data/ folder
beside the config; --fixtures path names it outright). Any Python 3 works:
3.11 and newer read the TOML with tomllib, older ones with the script’s
own small parser. The last line is the verdict:
claimed T1 T2, verified T1 T2
CI does the same on every push to main and fails if that line is anything else. The checker writes nothing to the portal apart from the audit row every CSV export leaves; its one write request is refused.
The isolation probe
The checker tries one peer probe and one participant probe. Passing it
means little on its own: a portal that returned 403 for every judge request
would pass. So
scripts/isolation_curl.sh
tries 94 things over HTTP, each with the exact status it must get back:
sh scripts/isolation_curl.sh # against http://localhost:8080
sh scripts/isolation_curl.sh http://localhost:9000 # or another base URL
It needs sh, curl and sed. First it reads the ids it needs with their
owners’ own tokens (one of judge A’s assignments, a project judge A
reviewed, the participant’s project, another team’s project), then:
| Group | Attempts | Expected |
|---|---|---|
| judge scores | judge B names judge A by fixture id and by email, and names the constant judge; the participant reads scores and lists assignments; no token; a made-up token; judge A reads their own | 403, 401, and 200 for the last |
| someone else’s review | judge B and the participant post scores to judge A’s assignment; judge B opens judge A’s scoresheet page | 404 |
| deadline | the participant submits and edits after the close; edits another team’s project; a judge and an anonymous caller create projects | 409, 403, 401 |
| exports | all eight kinds, anonymously, as the participant, as a judge, and as the organizer | 401, 403, 403, 200 |
| results before publication | anonymous, judge and participant read the fixture’s results | 404 |
| organizer pages | a judge’s and the participant’s tokens on every organizer page | 403 |
| community vote and comments | reading and casting ballots with no or a made-up token, voting where there’s no vote, a made-up voting link, unpublished results, commenting anonymously | 401, 404 |
| signed records, bundles and webhooks | judge B fetching judge A’s record or certificate, the participant and anonymous callers fetching records, exporting or importing bundles, opening the webhooks page; judge A fetching their own record | 403, 404, 401, and 200 for the last |
| public pages | the gallery, a submitted project and its comments, the API schema, the signing key, the embeddable gallery | 200 |
It prints one line per attempt and a total, and exits with status 1 if anything came back different:
== judge scores
ok 200 judge_a reads own scores
ok 403 judge_b names judge_a by fixture id
...
94 of 94 as expected
Two things to know when running it:
- It expects the fixture’s results to be unpublished. If you’ve
published them on the stack, the three results lines fail, correctly.
Reset with
docker compose down -v. - It speaks plain HTTP to port 8080, so it’s for the demo stack. A
deployment with
DJANGO_SECURE=1redirects it to HTTPS.
Like the checker, a passing run changes nothing: every write it tries is refused. The organizer’s exports leave their audit rows.
The test suite goes further than both (the same rules through the pages, the API, the admin and raw SQL against a real Postgres), but the probe is the one to run against a deployment, because it asks the running thing.
Build notes
A running log of what surprised me, what I changed my mind about, and why. Newest last.
Team names repeat in the fixture
I started with UNIQUE (event, name) on teams, which felt obviously right.
The seed fell over on its first run: the fixture has three different teams
called StillTrail and two each called AmberSwitch and OpenSignal. They have
different ids and different members, so they are different teams. Dropped
the constraint; teams are told apart by id, and the UI shows the id-based
link. Projects, by contrast, are checked for duplicates on purpose.
The planner’s first version wasn’t balanced
My first planner went project by project, giving each the least-loaded eligible judge. A property test with 12 projects, 6 judges and k=3 ended at loads 7, 7, 6, 6, 5, 5, not 6 each. Driving by judge instead (the idlest judge picks the neediest project) didn’t fix it either: by the end, the two idle judges already held both of the last projects that needed someone. Greedy can’t see that coming. What fixed it was a repair pass afterwards: move one of the plan’s new reviews from the busiest judge to the idlest one who is allowed to take it, until no move narrows the gap. Existing reviews are never moved. The property test now checks loads stay within one on 200 random events.
The confirmation link that confirmed itself
My first voting link confirmed the voter on GET, which is what my plan said and what every tutorial does. Then I remembered that corporate mail scanners (and some webmail previews) fetch every link in a message before the person sees it. With a single-use token, the scanner would use it up and the voter would click a dead link. Opening the link now shows a button, and the POST behind it confirms. One extra click, and the link survives being looked at.
A refused request that forgot it happened
The rate limiter counts hits in Postgres, inside the request’s transaction
(every request is atomic). My API views raised Conflict for an
over-budget ballot, DRF’s exception handler marks the transaction for
rollback, and the rollback took the rate-limit hit with it. So failed
requests were free, which is exactly backwards: the requests worth limiting
are the ones that fail. The API views now return their refusals as
responses instead of raising them, and the hit commits. The page views
already did.
Where the ballot’s privacy stops
I went back and forth on what vote.cast should record. The choices in
the clear would make the log a full record, but every organizer of the
event reads that log, and a community ballot is between the voter and the
tally. So the row has the voter, the credits spent, the number of projects
and a fingerprint of the ballot. A plain SHA-256 of something as small as a
quadratic ballot can be reversed by trying every ballot, so the fingerprint
is an HMAC keyed with the server’s secret. It still lets the abuse panel see
several voters casting the same ballot, which is the one thing it was for.
One inbox, one ballot, and not telling anyone who voted
My plan said a second sign-up with another spelling of the same inbox
(a.b+x@googlemail.com after ab@gmail.com) must be refused. Saying “that
inbox has already voted” would let anyone check whether a given address
voted. Instead the page reads the same either way, no second ballot is made,
the attempt is logged as vote.duplicate_refused for the abuse panel, and
a fresh link goes to the address just typed. (At first it went to the
address that signed up first; see the end of these notes for why not.) The
person who owns the inbox can still get in; nobody else learns anything.
Hiding results while the vote is open
Refusing to publish while voting is open wasn’t quite enough: an organizer could publish the judges’ results first and then open a vote, and voters would have the ranking in front of them. Published results now go back to a 404 for the public while a vote is open, and come back when it closes.
Tier T4: what surprised me
A route would have swallowed another. I first added
api/events/import at the end of urls.py, but the older
api/events/<slug:slug> matches first, with slug import, and would have
answered every import with 405. It sits above that pattern now, with a
comment saying why.
Pasted JSON can’t always be echoed back. The verify API returns the
record it was sent, so a record containing a lone surrogate ("\ud800",
perfectly legal JSON) made DRF’s renderer throw a 500. The malformed-input
test found it. Looking for its siblings turned up NaN, which Python’s
json reads but DRF’s strict output refuses. Both now fail the canonical
encoding step and come back as “not valid”, and both are test cases.
SET LOCAL outlives the function that sets it. The importer turns on
the deadline bypass for “its own transaction”, but inside a request with
ATOMIC_REQUESTS its transaction is only a savepoint: the bypass would have
stayed on for the rest of the request. Harmless for the seed, not for an
API endpoint, so the import now switches it off before returning (a failed
import’s savepoint rollback undoes it anyway).
The fixture importer trusted its file; a bundle importer can’t. Bundles
come from any organizer, and the fixture path wrote repo_url straight into
a link. A javascript: URL would have been stored XSS on the project page.
Bundles now go through the same URL, email and choice checks the forms use.
The round trip calibrates identically, not approximately. I expected to need a tolerance: floating point sums depend on order. But the export lists reviews by id and the import creates them in that order, so the fit sees the same sequence and every q matches to the last bit. The test asserts equality.
An iframe that reports scrollHeight can grow but never shrink:
document.documentElement.scrollHeight is never less than the iframe’s own
height. The embed reports the height of its content box instead; I checked
it in headless Chromium on a page that embeds the widget.
Three branches at once, and what the docs caught
I split the last stretch into three branches built side by side: the book, the public vote, and the T4 pieces. Two things I didn’t expect:
Both code branches took migration 0013. Each was right on its own branch and together they gave Django two leaf nodes. The webhooks migration became 0014 and depends on the voting one; nothing outside a scratch database had applied either, so renaming was safe. Next time I’d reserve numbers up front.
Writing the docs found real bugs. Explaining the calibration page line by
line turned up that the page and normalization_proof disagreed: the page
leaves out a project whose every reviewer the model ignores (prj_24, both
of whose judges are discordant), the proof didn’t, so JUDGING.md quoted a
ranking the product never shows. They also ran the agreement test with
different shuffle counts. Documenting the operations side found that
BALLOTBENCH_DEMO_SEED=0 still loaded the fixture event, that DJANGO_SECURE
made the health check fail on its own redirect, and that an organizer’s
?judge=jdg_24 could match a judge in another event, since fixture ids are
only unique per event. All fixed, each with a test. The lesson I keep
relearning: the fastest code review is trying to explain the code to a
stranger.
What a security review found in the vote and the duplicates
Void, read, unvoid. Organizers saw the tallies live and could undo a void, so voiding one ballot, reading the tallies and counting it again showed exactly what that voter chose, while every page said ballots were private. A void is now final, and while voting is open the organizer page shows how many ballots there are, not the tallies. After the close the tallies are there, and a void still subtracts a ballot you can see; but it is gone for good, so reading it costs the organizer that vote.
The first spelling of an inbox owned its links. Asking for a link as
alice+nope@yahoo.com sent every later link for alice@yahoo.com to the
+nope address. For providers where a +tag is its own inbox, that’s a
stolen ballot. The link now goes to whatever was typed; a new link replaces
the old one. Quoted local parts ("a.b"@gmail.com) are refused: legal, but
nobody needs one to vote, and normalizing them right is a rabbit hole.
normalize_email also used to raise on +x@example.com, and because it runs
over every team member’s address, one such account broke every email
voter’s ballot page. It no longer raises.
Duplicates by pk. The detector ordered projects by pk, drafts included,
and every submit re-flagged the whole event. A draft made early and filled
in later with another team’s public title and repo made the real project
the “duplicate”, taking it out of judging and off the ballot; and an
organizer’s “it’s different” lasted until the next submit. Now only
submitted projects count, earlier means submitted_at, a submit only flags
the project being submitted, and duplicate_cleared keeps an organizer’s
call.
What a stranger found in an hour
With every test green, I had someone follow the book’s tour from a fresh clone, as a judge would, and write down everything that disagreed with it. It found 28 things. The ones that stung:
- “Save draft” after submitting quietly replaced the submitted scores. The review stayed “submitted”, with the old timestamp, and calibration read the new numbers. Every test submitted once and stopped. A submitted review now only resubmits.
- Deleting a rubric criterion after scoring was a 500. Changing a weight was refused politely by the trigger; deleting went through a foreign key set to RESTRICT, and nobody had wrapped that path. The same shape as the track bug the security review found an hour earlier: a database refusal reached from a path the app didn’t guard.
- The venue problem. Rate limits per address are right against a flood and wrong at a hackathon, where three hundred people share one NAT address. The walkthrough locked itself out after twenty ordinary logins. Per-address limits now scale for a crowd, and a per-account limit does the job of stopping someone guessing one person’s password.
- The offline proof published no port.
internal: truenetworks can’t publish ports, so following the README gave a stranger nothing to open. The honest instruction is simpler: switch off the Wi-Fi.
None of these were security holes and all of them would have been on camera in a demo. Reading the docs aloud against the running thing is a test nothing else replaces.