Build Rundown & Issue Map
Written 2026-09-25 · AI at Work in IO Pipeline
Spark Monitor — GPU 96%, 105.7 / 119.6 GiB unified memory, 72 tok/s
GX10-9BB4 · up 16d
| Component | Specification |
| System |
ASUS Ascent GX10 (NVIDIA DGX Spark, GB10, 20-core Arm CPU, 128GB unified memory, aarch64) |
| OS |
Linux / DGX OS |
| Storage |
4TB NVMe at /mnt/ai_data |
| Access |
Headless, via NoMachine |
| Model |
Qwen3.6-35B-A3B-NVFP4 |
What this is
A daily research brief for HR and organizational leaders, produced by a local AI model with one paid call a week for the part worth paying for.
The economics are the design constraint. The daily loop runs on hardware already bought, at no marginal cost. The weekly proposal call costs about three cents.
What runs now
| Step | Schedule | Model call | Output |
| Collect |
07:15 weekdays |
None |
Inbox items with IDs |
| Generate queries |
07:20 |
Local AI · 1 call |
Web search queries (dynamic) |
| Search |
07:21 |
None |
DuckDuckGo results → inbox |
| Filter |
07:30 |
None |
I/O relevance filter applied |
| Rate |
07:35 |
Local AI · 1 per 8 items |
Ratings file |
| Brief |
07:55 |
Local AI · 1 call |
Daily brief → Telegram |
| Weekly review |
Manual |
None / ~3¢ |
Focus question for the week |
Twelve scripts, each writing a file the next one reads. A failure stops at one step instead of corrupting the next.
The scripts
| Script | Role | Lines |
| collect.py |
Polls RSS feeds, downloads new items |
1,806 |
| gen_queries.py |
Generates web search queries from gaps in recent coverage |
272 |
| search.py |
Runs DuckDuckGo searches with the generated queries |
404 |
| focus.py |
Chooses the week's editorial focus question |
960 |
| rate-items.py |
Rates each item against the rubric |
527 |
| focus-filter.py |
Filters items by I/O psychology relevance |
141 |
| brief-lint.py |
Verifies brief output — 31 error conditions, 17 warnings |
571 |
| feed-health.py |
Monitors feed connectivity and silent sources |
336 |
| engine-probe.py |
Health-checks the local vLLM model |
271 |
| engine-stress.py |
Stress-tests vLLM under load |
206 |
| calibrate-migrate.py |
Calibrates ratings across model versions |
184 |
| verdict.py |
Records final verdict on rated items |
209 |
The Issue Map
Grouped by failure class. Each class produced one structural change rather than another instruction.
Class A The local model doesn't work
Three days were lost before any pipeline work. Each symptom was a misconfiguration, not a capability gap.
| Symptom | Actual cause | Fix |
| Model hangs, returns nothing |
Model was spending all tokens on reasoning. The chat template was enabling thinking mode. |
Force disable thinking in chat template. Patched and committed. |
| Engine wedges mid-run |
CUDA graph replay crashed while the API server kept returning success codes. |
Disable async scheduling. Container exits and restarts cleanly. |
| Over 130 items triaged by mistake |
A delegation tool was enabled on the interactive path and spawned parallel workers. |
Exclude it from cron; same treatment as CLI toolset. |
Lesson: only a real completion is a valid health check. A status code of 200 proves nothing.
Class B Asking the model to hold a fact a script already had
The load-bearing class. Every defect worth fixing was an instance of this. The model was being asked to compose facts when the scripts already knew them.
| The fact | What went wrong | Where it lives now |
| Item IDs |
A skill told the model to invent its own IDs |
Collector assigns; model copies |
| Ratings |
The brief re-derived its own ratings and rated upward — promoting weak items and burying strong ones |
Python reads the ratings file directly. No model call needed. |
| Age in days |
Wrong on four items while publication dates were correct |
Print the date; let the reader subtract |
| Qualifying count |
The model miscounted its own output |
Script counts after generation |
The rule: A fact a script holds, the script writes. The model copies forms, never composes facts.
Class C Contradictory instructions
Two documents said different things. The model obeyed one, was blamed for the result, and the cycle repeated.
| Conflict | Consequence |
| Skill said mint IDs; user instruction said copy them |
It obeyed the user — by luck, not design |
| Doc said "weekly brief"; system produces a daily one |
Loaded wrong context into every run |
| Template prose leaked into output |
Raw instructions like "at most one, omit if none" appeared in delivered briefs |
Fix: Bump the skill version, strip stale instructions, and provide filled examples the model can copy verbatim.
Class D Declared inputs, nothing wired up
Four files were referenced as though in use. No script ever sent them to any model call.
- Focus file — the week's editorial question was never included. The brief invented its own and ignored relevant papers.
- Audience file — given to one script but not the other. The two components disagreed because one of them didn't know who the blog was for.
- Evidence types — eight controlled labels, never sent. The brief wrote the rating in the evidence slot instead.
- Covered file — empty. No script appends to it.
Class E Checks that trusted their subject
The subtlest class. The first fix read the ratings path from the brief's own pointer line — written by the model. A check taking its reference from the text under check verifies nothing.
Fix: A verifier must not take its reference from its subject. Where it cannot get an independent reference, it should report that it could not check.
Class F Supply, diagnosed by evidence
The inbox went thin and there were three plausible stories: a cap, dedup suppression, or genuinely quiet feeds. Each was ruled in or out by evidence:
- Empty holding list → ruled out the cap
- Seen list at 32 entries → ruled out suppression
A 14-day dry run returning 87 candidates from 18 feeds → proved it was a two-day lookback against a collector younger than the archive |
How it actually proceeded
Sep 13 — Skills, v2
2 packages — scaffolding the daily loop
Sep 14 — Brief linter
1 package — 145 lines. The linter earns its place in its first round.
Sep 16 — Collector
1 package — the chain becomes deterministic
Sep 18 — Rate items, rubric v7
2 packages — judgement moves to its own script
Sep 19 — Item IDs
1 package — IDs stop being the model's job
Sep 20 — Focus ×4, ratings flow, brief appraise, polish, calibration, feed health, placeholder, leak, stamp
12 packages — twelve in one day. The brief stops re-rating; appraisal fields land; three prose-leak fixes.
Sep 23 — Dynamic queries
2 packages — gen_queries.py and search.py add DuckDuckGo to the pipeline; items seen grows from 32 to 384+
Sep 23 — Custom queries file
1 package — custom-queries.txt lets Lee inject keywords that influence search
Sep 23 — Search.py
1 package — actual web search replaces the gap between collection and rating
Sep 25 — Blog post published
1 post — "AI Legal Costs" goes live on ai-at-work.io; pipeline reaches 539 seen items, 12 scripts
The clustering on Sep 20 is not thrash. It is what happened once the root cause was understood as structural rather than instructional: each fix exposed the next thing the model was still being asked to compose.
brief-lint.py: 145 lines → 571 lines, 31 error conditions and 17 warnings. The pipeline itself is 5,887 lines across 12 scripts.
The one rule
Everything compresses to a single sentence, now the first hard rule in the project docs:
Facts the scripts hold, the scripts write. Ids, ratings, the count, the rubric version — all script-written, copied by the model, never composed by it.
Still open
-
arXiv returns 406 — curl probes drafted, never run
open
-
HBR TLS refusal — publisher refusing non-browser clients
open
-
Moonshots feed works (HTTP 200, 297 entries), but the focus filter blocks it
investigating
-
covered.md — placeholder with header, no actual entries yet
partial
-
The weekly post — picks a question but nothing drafts the post
partial