USYD CODING FEST 2026·PROJECT 18·UG SENIOR

From CT scan to signed report.

Point-prompt segmentation, four specialist AI agents and a QC gate draft a structured, evidence-cited CT report — any body region. The radiologist reviews, revises and approves every word. AI drafts. Doctors decide.

REAL MODEL OUTPUT · SAM-MED2D · DEMO CT MASK RETURNED · 5–17 MS
4+1
specialist agents
5–17 ms
mask decode
26–55 ms
evidence retrieval

BUILT AND DEPLOYED ON

01THE FILM

Ninety seconds. Nothing staged.

The booth film is rendered from what the product actually does: the CT slices come from the demo backend, every mask is real SAM-Med2D inference, and the latencies on screen are measured — not invented.

MASKS — REAL INFERENCE (POINT IoU 0.725 · BOX 0.805 · DECODER 5–17 MS MEASURED). FULL FILM PLAYS WITH VOICE-OVER; CHAPTERS ARE SILENT BY DESIGN.

02THE PROBLEM

Radiology needs speed and consistency. Isolated AI delivers neither.

CT interpretation still varies between clinicians, and most AI tools bolt on as black boxes — outputs that are hard to verify, hard to trace, and hard to fit into how radiologists actually work.

ISOLATED AI TOOLS — THE STATUS QUO

Single-shot outputs with no evidence trail — where did this finding come from?
Verification is manual: the clinician re-does the work before trusting it.
No place in the reporting workflow — results live outside the system of record.

SOMA — CLINICIAN-CENTERED BY DESIGN

One workflow: segment → findings → diagnosis → draft → evidence → review.
Every claim cites its source — Lung-RADS, ICD-10, Fleischner/ACR and RadLex built in.
Human-in-the-loop is the point, not a checkbox: agents draft, the radiologist signs.

SOMA is not built to replace physicians — it is built so physicians can trust what AI drafts.

03THE SYSTEM

Five specialists. One state machine. A doctor at the end.

An Orchestrator drives every case through a finite state machine. Agents hand off in sequence, the knowledge base grounds each phase — and only a human moves a case past Draft Ready.

CREATED ANALYZING DRAFT READY REVISING APPROVED
  1. PHASE 1Radiologist Location · size · margins · morphology. SONNET 5 · VISION
  2. PHASE 2Pathologist Ranked differential → working diagnosis. SONNET 5
  3. PHASE 3Report Writer ACR-structured draft, citations attached. SONNET 5
  4. PHASE 4QC Reviewer Consistency · completeness · citations. SONNET 5
AlignmentREVISION ONLY Routes each doctor request to the one specialist it actually needs. SONNET 5 · ON DEMAND

Revision is a conversation — and a router. The Alignment Agent maps 22 request intents onto 5 handlers. Only the target agent re-executes; downstream stages re-validate instead of regenerating. Nothing changes silently.

  • “soften the wording”→ REPORT WRITER
  • “re-measure the lesion”→ RADIOLOGIST
  • “question the diagnosis”→ PATHOLOGIST
  • “this contradicts itself”→ QC REVIEWER
  • “fix the typo”→ NO AGENT RE-RUN

04MEASURED, NOT CLAIMED

The numbers we can defend at the booth.

441
curated guideline entries
LUNG-RADS · ICD-10 · FLEISCHNER/ACR · RADLEX
768-d
PubMedBERT embeddings
PGVECTOR · HNSW INDEX · POSTGRESQL
26–55 ms
evidence retrieval
MEASURED PER CALL, IN PRODUCTION PATH
0.92
IoU — point vs box prompt
SAME LESION BOUNDARY EITHER WAY
270 ms
self-hosted intent classification
FINE-TUNED QWEN3-1.7B · DEPLOYED ON AWS
100%
mode & intent accuracy
VS ~95% / ~85% ON GPT-4O-MINI

05DATA RESIDENCY · COMPLIANCE

Patient data never leaves Australia.

All LLM inference runs on Anthropic Claude via AWS Bedrock — pinned to Australian regions. Compute sits next to the data in Sydney, and no API key sits on the server.

au. inference profileBedrock keeps every model call inside Australian AWS regions — enforced by the profile, not by promise.
Sydney, end to endBackend on AWS EC2 and Neon PostgreSQL both in ap-southeast-2 — compute sits next to the data.
Zero static keysIAM instance-profile auth only — no long-lived API keys on the server to leak.
Everything on AWSData, models and our own fine-tuned weights all deploy on the same AWS stack — nothing runs on a laptop.

GROUNDING — LUNG-RADS v2022 · ICD-10 · FLEISCHNER/ACR · RADLEX — CURATED ONCE, RETRIEVED BY EVERY AGENT AT EVERY PHASE.

06ENGINEERING DEEP-DIVE

We found our bottleneck — and out-engineered the API.

In the revision loop, every doctor message must be understood before anything runs: revise, re-analyse, or just a question — and for which agent. The default — GPT-4o-mini — was metered, ~500 ms and not accurate enough. So the team fine-tuned Qwen3-1.7B on 5,000 samples with one RTX 4090, and replaced the API call outright.

Qwen3-1.7B — the intent router 270 MS · ON AWS

It reads every chat message and decides three things at once: what the doctor wants (mode), which agent owns it (target), and how much re-runs (scope). Three real messages, three different routes:

“Commit to Lung-RADS 4A and tighten the follow-up.” MODE · REVISEReport WriterRE-RUNS · IMPRESSION ONLY
“Re-measure the nodule against the prior CT.” MODE · RE-ANALYSERadiologistRE-RUNS · FINDINGS FLOW DOWNSTREAM
“Why a 3-month interval?” MODE · QUESTIONAnswered in chatCITED · NOTHING RE-RUNS
METRICGPT-4O-MINIQWEN3-1.7B · FINE-TUNED
mode accuracy~95%100%
intent accuracy~85%100%
latency~500 ms270 ms · 1.85× faster
cost / 1K tokens$0.15$0
LLAMA-FACTORY RTX 4090 × 1 5,000 SAMPLES · ALPACA 2 EPOCHS EVAL LOSS 0.07 TRAINED IN 9.6 MIN

This is the SOMA pattern: measure the bottleneck, train the smallest model that wins on every axis, deploy it on our own AWS stack.

07THE ARTIFACT

Read a report SOMA drafts.

A representative chest-CT draft in the exact structure the pipeline emits — rendered inline, no backend. If the live demo ever goes down at the booth, this page is the demo.

SOMA · MEDICAL REPORTDRAFT
CTCHEST1.0 MM AXIALSLICE 042/120GENERATED 2026-07-21 09:42 AEST

Clinical Context

58-year-old male, current smoker, 32 pack-years. Baseline low-dose CT lung-cancer screening. No prior chest imaging available for comparison.

Technique

Low-dose non-contrast chest CT, 1.0 mm axial reconstructions. ROI segmentation with SAM-Med2D — 2 positive / 1 negative point prompts on slice 042/120.

Findings

Solid, well-circumscribed nodule in the right upper lobe (segmented ROI, Mask 1), measuring 9 × 8 mm. Margins are smooth without spiculation, cavitation or calcification. RadLex · RID3875

No suspicious mediastinal or hilar lymphadenopathy. No pleural effusion or consolidation. The remaining parenchyma is clear. ICD-10 · R91.1

Impression

  1. Solid 9 mm right-upper-lobe nodule at baseline screening — Lung-RADS category 4A. Lung-RADS v2022 · 4A
  2. No other actionable findings.
4A Lung-RADS 4A — suspicious. Estimated 5–15% probability of malignancy. Management: LDCT at 3 months; PET/CT may be considered for solid nodules ≥ 8 mm.

Recommendation

Follow-up low-dose chest CT in 3 months. For a solid nodule of this size, short-interval surveillance is concordant with Lung-RADS 4A management and Fleischner Society guidance. Fleischner 2017 Lung-RADS v2022

QC REVIEW — 4 AUTOMATED CHECKS

Measurement consistent across Findings and Impression (9 mm).
All guideline citations resolved in the knowledge base (4/4).
Lung-RADS category consistent with nodule size and baseline context.
needs_clinical_correlation — correlate with smoking history and any prior malignancy before sign-off.
DRAFT

Awaiting radiologist review. SOMA drafts; the physician revises and approves. Nothing is final until signed.

Illustrative sample with synthetic patient data, generated for demonstration. SOMA is a research prototype for decision support — not a registered medical device. Every output requires review and approval by a qualified clinician.

08NEXT · TEAM

What we build next — and who builds it.

  • N1ACR-compliant templates, dual outputPaired clinician and patient versions of every report.
  • N2Per-edit diff reviewAccept / reject controls on every single change, GitHub-style.
  • N3Production hardeningStronger authentication, rate limiting and schema validation across the API surface.
  • N4Beyond CTMore modalities — MRI, X-ray, ultrasound — and larger clinical datasets.
  • T1
    Ruixuan Liao
    USYD · PROJECT 18
  • T2
    Mingzhe Cai
    UNSW · PROJECT 18
  • T3
    Siyu Chen
    USYD · PROJECT 18
  • T4
    Kening Hu
    USYD · PROJECT 18

BOOTH 18 · USYD CODING FEST 2026

Give it a case.

The live demo runs on the real pipeline — synthetic data only. Or read the sample report above without signing in.