Build Rundown & Issue Map

Written 2026-09-25 · AI at Work in IO Pipeline

AI at Work pipeline architecture diagram
Pipeline architecture — sources → collect → rate → brief Created by Nano Banana · 2026-09-25
Spark Monitor showing system running Qwen3.6-35B-A3B-NVFP4 — GPU 96%, 105.7/119.6 GiB unified memory, 72 tok/s decode throughput
Spark Monitor — GPU 96%, 105.7 / 119.6 GiB unified memory, 72 tok/s GX10-9BB4 · up 16d
ComponentSpecification
System ASUS Ascent GX10 (NVIDIA DGX Spark, GB10, 20-core Arm CPU, 128GB unified memory, aarch64)
OS Linux / DGX OS
Storage 4TB NVMe at /mnt/ai_data
Access Headless, via NoMachine
Model Qwen3.6-35B-A3B-NVFP4
22
Build packages
5,887
Lines of code
539
Items seen
What this is

A daily research brief for HR and organizational leaders, produced by a local AI model with one paid call a week for the part worth paying for.

The economics are the design constraint. The daily loop runs on hardware already bought, at no marginal cost. The weekly proposal call costs about three cents.

What runs now
StepScheduleModel callOutput
Collect 07:15 weekdays None Inbox items with IDs
Generate queries 07:20 Local AI · 1 call Web search queries (dynamic)
Search 07:21 None DuckDuckGo results → inbox
Filter 07:30 None I/O relevance filter applied
Rate 07:35 Local AI · 1 per 8 items Ratings file
Brief 07:55 Local AI · 1 call Daily brief → Telegram
Weekly review Manual None / ~3¢ Focus question for the week

Twelve scripts, each writing a file the next one reads. A failure stops at one step instead of corrupting the next.

The scripts
ScriptRoleLines
collect.py Polls RSS feeds, downloads new items 1,806
gen_queries.py Generates web search queries from gaps in recent coverage 272
search.py Runs DuckDuckGo searches with the generated queries 404
focus.py Chooses the week's editorial focus question 960
rate-items.py Rates each item against the rubric 527
focus-filter.py Filters items by I/O psychology relevance 141
brief-lint.py Verifies brief output — 31 error conditions, 17 warnings 571
feed-health.py Monitors feed connectivity and silent sources 336
engine-probe.py Health-checks the local vLLM model 271
engine-stress.py Stress-tests vLLM under load 206
calibrate-migrate.py Calibrates ratings across model versions 184
verdict.py Records final verdict on rated items 209
The Issue Map

Grouped by failure class. Each class produced one structural change rather than another instruction.

Class A The local model doesn't work

Three days were lost before any pipeline work. Each symptom was a misconfiguration, not a capability gap.

SymptomActual causeFix
Model hangs, returns nothing Model was spending all tokens on reasoning. The chat template was enabling thinking mode. Force disable thinking in chat template. Patched and committed.
Engine wedges mid-run CUDA graph replay crashed while the API server kept returning success codes. Disable async scheduling. Container exits and restarts cleanly.
Over 130 items triaged by mistake A delegation tool was enabled on the interactive path and spawned parallel workers. Exclude it from cron; same treatment as CLI toolset.
Lesson: only a real completion is a valid health check. A status code of 200 proves nothing.

Class B Asking the model to hold a fact a script already had

The load-bearing class. Every defect worth fixing was an instance of this. The model was being asked to compose facts when the scripts already knew them.

The factWhat went wrongWhere it lives now
Item IDs A skill told the model to invent its own IDs Collector assigns; model copies
Ratings The brief re-derived its own ratings and rated upward — promoting weak items and burying strong ones Python reads the ratings file directly. No model call needed.
Age in days Wrong on four items while publication dates were correct Print the date; let the reader subtract
Qualifying count The model miscounted its own output Script counts after generation
The rule: A fact a script holds, the script writes. The model copies forms, never composes facts.

Class C Contradictory instructions

Two documents said different things. The model obeyed one, was blamed for the result, and the cycle repeated.

ConflictConsequence
Skill said mint IDs; user instruction said copy them It obeyed the user — by luck, not design
Doc said "weekly brief"; system produces a daily one Loaded wrong context into every run
Template prose leaked into output Raw instructions like "at most one, omit if none" appeared in delivered briefs

Fix: Bump the skill version, strip stale instructions, and provide filled examples the model can copy verbatim.

Class D Declared inputs, nothing wired up

Four files were referenced as though in use. No script ever sent them to any model call.

Class E Checks that trusted their subject

The subtlest class. The first fix read the ratings path from the brief's own pointer line — written by the model. A check taking its reference from the text under check verifies nothing.

Fix: A verifier must not take its reference from its subject. Where it cannot get an independent reference, it should report that it could not check.

Class F Supply, diagnosed by evidence

The inbox went thin and there were three plausible stories: a cap, dedup suppression, or genuinely quiet feeds. Each was ruled in or out by evidence:

How it actually proceeded
Sep 13 — Skills, v2
2 packages — scaffolding the daily loop
Sep 14 — Brief linter
1 package — 145 lines. The linter earns its place in its first round.
Sep 16 — Collector
1 package — the chain becomes deterministic
Sep 18 — Rate items, rubric v7
2 packages — judgement moves to its own script
Sep 19 — Item IDs
1 package — IDs stop being the model's job
Sep 20 — Focus ×4, ratings flow, brief appraise, polish, calibration, feed health, placeholder, leak, stamp
12 packages — twelve in one day. The brief stops re-rating; appraisal fields land; three prose-leak fixes.
Sep 23 — Dynamic queries
2 packages — gen_queries.py and search.py add DuckDuckGo to the pipeline; items seen grows from 32 to 384+
Sep 23 — Custom queries file
1 package — custom-queries.txt lets Lee inject keywords that influence search
Sep 23 — Search.py
1 package — actual web search replaces the gap between collection and rating
Sep 25 — Blog post published
1 post — "AI Legal Costs" goes live on ai-at-work.io; pipeline reaches 539 seen items, 12 scripts

The clustering on Sep 20 is not thrash. It is what happened once the root cause was understood as structural rather than instructional: each fix exposed the next thing the model was still being asked to compose.

brief-lint.py: 145 lines → 571 lines, 31 error conditions and 17 warnings. The pipeline itself is 5,887 lines across 12 scripts.

The one rule

Everything compresses to a single sentence, now the first hard rule in the project docs:

Facts the scripts hold, the scripts write. Ids, ratings, the count, the rubric version — all script-written, copied by the model, never composed by it.
Still open