Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

The acceptance checker

Dogfood’s organizers give every team the same checker, run.py: a standard-library Python script that reads a team’s .dogfood.toml, makes seven HTTP requests against the running portal, and prints a report of which tiers it could verify. This chapter explains the contract as ballotbench meets it, and the stricter probe I wrote to go past it.

.dogfood.toml, line by line

[portal]
base_url = "http://localhost:8080"

[tiers]
claimed = ["T1", "T2"]
pitch = "A self-hosted hackathon portal whose judging you can defend: weighted rubrics, calibrated judges, and deadlines and isolation enforced in the database."

[auth]
organizer   = "Authorization: Bearer bb_demo_organizer_5c1e0a"
judge_a     = "Authorization: Bearer bb_demo_judge_a_8d24f1"
judge_b     = "Authorization: Bearer bb_demo_judge_b_3a9e77"
participant = "Authorization: Bearer bb_demo_participant_61b0c4"

[routes]
gallery      = "/projects"
submit       = "/api/events/sample-hack-2026/projects"
judge_scores = "/api/judge/scores"
peer_scores  = "/api/judge/scores?judge=jdg_24"
csv_export   = "/api/events/sample-hack-2026/export/scores.csv"
LineMeaning
base_urlwhere the checker sends every request: the port docker compose up publishes
claimedthe tiers this entry claims: T1 (submissions and the gallery) and T2 (judging and isolation), the ones run.py has checks for. The public vote (T3) and the stretch pieces (T4) are built and tested, but the checker can’t verify them, so they aren’t claimed.
pitcha one-line description of the entry; run.py doesn’t read it
organizer, judge_a, judge_b, participanta complete HTTP header for each role. The checker splits it at the first : and sends it as is. These are the demo tokens the seed creates when BALLOTBENCH_DEMO_SEED=1.
gallerythe public page listing submitted projects
submitwhere a participant creates a project, here in the fixture event, which is closed
judge_scoreswhere a judge reads their own scores
peer_scoresthe URL that would return judge A’s scores: jdg_24 is judge A’s id in the fixture
csv_exportan organizer’s CSV export

The four accounts are chosen so the checks mean something: judge_a (jdg_24) and judge_b (jdg_29) both have fixture reviews and share no project, so each has scores the other must not see, and the participant is on a real fixture team (NorthKiln), so the late submission is refused for being late, not for having no team.

The seven checks

TierCheckThe requestPasses whenWhat answers it
T1gallery is publicGET /projects, no header200the gallery view, which lists submitted projects to anyone
T1project from fixtures shownthe same responseit contains the title of one of the fixture’s first three projectsthe seed imports the fixture on boot, and the gallery orders page one by id, so Glass Signal, Small Meadow and Deep Compass lead
T1closed event refuses submissionsPOST to submit as participant, with a title and summaryany 4xx409 with the close time: the API checks the window on the database’s clock, and the project_deadline trigger would refuse it anyway
T2judge sees own scoresGET judge_scores as judge_a200/api/judge/scores returns the caller’s own reviews
T2judge cannot see peer scoresGET peer_scores as judge_b401 or 403403: a judge may only name themselves, and gets a refusal, never an empty list
T2participant blockedGET judge_scores as participant401 or 403403 judges only
T2csv export worksGET csv_export as organizer200, and the first line has a commathe scores export, whose first line is the CSV header

The checker is lenient in places (any 4xx for the late submission, 401 or 403 for the peer probe). ballotbench answers with the specific code in each case; the API chapter has the contract.

A tier counts as verified only if every one of its checks passes and every tier below it is verified too. run.py has checks for T1 and T2 only, so those are the most it can verify.

Regenerating acceptance-report.txt

The report in the repository is the checker’s output against the Docker build on a fresh volume. To make it again, from the repository root:

docker compose down -v
docker compose up -d --build --wait
python3 run.py .dogfood.toml > acceptance-report.txt

Run it from the root so it finds fixtures.json (it looks in the current directory, next to run.py, next to the config, and in a data/ folder beside the config; --fixtures path names it outright). Any Python 3 works: 3.11 and newer read the TOML with tomllib, older ones with the script’s own small parser. The last line is the verdict:

claimed T1 T2, verified T1 T2

CI does the same on every push to main and fails if that line is anything else. The checker writes nothing to the portal apart from the audit row every CSV export leaves; its one write request is refused.

The isolation probe

The checker tries one peer probe and one participant probe. Passing it means little on its own: a portal that returned 403 for every judge request would pass. So scripts/isolation_curl.sh tries 94 things over HTTP, each with the exact status it must get back:

sh scripts/isolation_curl.sh                         # against http://localhost:8080
sh scripts/isolation_curl.sh http://localhost:9000   # or another base URL

It needs sh, curl and sed. First it reads the ids it needs with their owners’ own tokens (one of judge A’s assignments, a project judge A reviewed, the participant’s project, another team’s project), then:

GroupAttemptsExpected
judge scoresjudge B names judge A by fixture id and by email, and names the constant judge; the participant reads scores and lists assignments; no token; a made-up token; judge A reads their own403, 401, and 200 for the last
someone else’s reviewjudge B and the participant post scores to judge A’s assignment; judge B opens judge A’s scoresheet page404
deadlinethe participant submits and edits after the close; edits another team’s project; a judge and an anonymous caller create projects409, 403, 401
exportsall eight kinds, anonymously, as the participant, as a judge, and as the organizer401, 403, 403, 200
results before publicationanonymous, judge and participant read the fixture’s results404
organizer pagesa judge’s and the participant’s tokens on every organizer page403
community vote and commentsreading and casting ballots with no or a made-up token, voting where there’s no vote, a made-up voting link, unpublished results, commenting anonymously401, 404
signed records, bundles and webhooksjudge B fetching judge A’s record or certificate, the participant and anonymous callers fetching records, exporting or importing bundles, opening the webhooks page; judge A fetching their own record403, 404, 401, and 200 for the last
public pagesthe gallery, a submitted project and its comments, the API schema, the signing key, the embeddable gallery200

It prints one line per attempt and a total, and exits with status 1 if anything came back different:

== judge scores
ok   200 judge_a reads own scores
ok   403 judge_b names judge_a by fixture id
...

94 of 94 as expected

Two things to know when running it:

  • It expects the fixture’s results to be unpublished. If you’ve published them on the stack, the three results lines fail, correctly. Reset with docker compose down -v.
  • It speaks plain HTTP to port 8080, so it’s for the demo stack. A deployment with DJANGO_SECURE=1 redirects it to HTTPS.

Like the checker, a passing run changes nothing: every write it tries is refused. The organizer’s exports leave their audit rows.

The test suite goes further than both (the same rules through the pages, the API, the admin and raw SQL against a real Postgres), but the probe is the one to run against a deployment, because it asks the running thing.