6 min read

A team of Claude agents designed 53 toys for my 3D printer

Claude agents turned a short brief into 53 toy ideas for my 3D printer and built a page you can play with for each of my 16 finalists. Two of them in detail, and a file you can hand to your own agent to run the same process.

My 6-year-old walks across the living room holding a stone that gets warmer as she gets closer to the treasure. The stone doesn’t exist yet. A team of Claude agents designed it for my 3D printer, along with 52 other toys, and it made my final 16. Below is how the agents worked and a file you can give your own agent to run the same process.

Drawn by
a team of Claude agents
Checked by
a fresh agent
Picked by
me, on my phone

15–25 Sep

A short brief and a loop

My kids have enough toys. My 3D printer doesn’t have enough work. Claude’s agents were happy to help with the second problem and ignore the first. I asked them for a toy that is unique in some way and combines the 3D printer, electronics and AI, to see whether the ideas can come from AI too, and how close they get to something I would actually build.

Claude Code is an AI assistant that runs on my computer and can start other Claude agents. My main session wrote a brief for each one, and each agent starts like a new chat, with none of the conversation that came before. The briefs grew into a loop where agents invent toys, a fresh agent that invented none of them checks the result, and the next round fixes what it found. The toys I like best then get a page with a demo you can play before anything is printed.

The toys are for my kids, so these rules ran through the work, some from my briefs and some the inventors set for themselves:

  • The Gemini API and ElevenLabs didn’t allow users under 18 in their terms when the agents checked in September. Anthropic allows products for minors with safeguards. Meshy (text to 3D) requires users to be 14 or older with a parent’s consent, so a parent operates it. Terms change, so your own agent should check them again.
  • A child’s voice never leaves the house. Speech recognition (whisper.cpp) and Czech text-to-speech (Piper) run locally.
  • Push-to-talk only. The toy never listens on its own.
  • Code holds the game state and the rules. The AI talks and picks from a safe list, and its output is validated before it moves a servo or opens a lock. One of the toys is a guardian that players try to sweet-talk out of a password, and only code can open its chest.
  • No loose coin cells, magnets sealed inside the print, and nothing small enough to swallow in anything the 6-year-old plays with.

Two of my favorites are built on the rule that code holds the state

Runes Say a spell tonight and hold it as a printed rune tomorrow. Ages 6–12 see the toy →

Where the spell goes

  1. A child tells the Book what the spell should do.
  2. Claude turns the sentence into a rule from a safe list.
  3. The Book reads the rule back.
  4. Overnight the rule goes into the rune’s chip.

Saved spells run with no network and no model.

Vault Guardian Talk the guardian out of his password, or wait till he dozes. Ages 12+ see the toy →

Who can open the chest

The players talk to the guardian, a language model. Only the game code is wired to the lock. There is no wire from the model to the lock.playersthe modelgame codethe locktalkopensno tool for it

24 Sep

First try: two agents with the same homework

One research agent collected 44 examples of AI meeting physical objects, a second covered printable mechanisms and toy safety, and a third scraped two Czech electronics shops for 68 parts with prices. Two inventors I had asked for, a free-thinking Dreamer and a grounded Engineer, then got the same brief and the same research and wrote 20 toys each, down to a sample dialogue and a priced parts list.

My main session picked a shortlist, and four agents built a page with a playable demo for each. By the evening round one had 43 toys (seven of them taken straight from the research), which is more than I can read on a phone.

25–28 Sep

A fresh agent finds the same toy twice

The next morning a fresh agent got round one’s inputs and outputs, without my main session’s notes, and was asked to diagnose it and propose a better process.

What the fresh agent counted

  • 42 of 43 need electronics and a running Mac

  • 32 of 43 use the AI mainly for words, spoken or printed

  • 3 of 43 are played outdoors

One square per toy.

My brief had already handed both inventors one skeleton: a cheap ESP32 board in the toy, a small server on the Mac, Claude behind it. On top of that they picked 4 of the same 5 starting points from the research, gave three ideas the same name and three more the same core, so 40 ideas came to about 34 distinct ones. Each inventor had also graded its own homework. Music and rhythm were missing as a kind of play, and I play the drums. The one toy I had designed and printed this year is a table hockey set with no electronics at all, the kind of toy round one never produced.

The next round had two new inventors with different reading lists, and an agent that invented nothing did the scoring. I added ten of that round’s toys, which took the list to 53. I went through all 53 on my phone anyway and kept 16. Fourteen of them still come from round one.

24–28 Sep

Play it before it exists

Each of the 16 has its own page with a playable demo at toys.malacka.cz, and every toy got its own visual world. I wanted a parent or maker who opens a page to think “I want this”.

Hot and Cold for Real is the stone from the top of this post. A parent hides three chests with a Bluetooth beacon inside each and photographs the spots. Claude looks at the photos and writes the riddles. On the page you drag the stone through a flat you’ve never been to. It reads the beacons’ jumpy signal the way a real stone would, and its palm plate warms up as you get closer. As the page admits, temperature changes over seconds, which is too slow to steer by, so the first version stays cold. It guides with vibration, light and voice, and the warmth waits until the game catches on at home.

Runes is the one I want to build most. A child tells the Book, a book-shaped stand, what a spell should do. Claude turns the sentence into a small rule, the Book reads it back, and the 3D printer makes a rune with that rule stored in an NFC chip sealed inside. On the page, the little sister gives her big brother a rune. He sets it on his own Book, circles his wand, and his room turns her shade of blue. Saved spells run offline with no AI, at a friend’s house too, as long as the friend has a Book and a wand. The room lights answer only at home. The runes on the page are 30 mm across, which toy-safety rules count as small enough to swallow, so the build plan makes them 40.

Two of the 16 pages, recorded in a browser

Hot and Cold for Real: the stone reads the beacons as you walk toward the first chest.Open the page
Runes: big brother casts his sister’s spell and his room turns her shade of blue.Open the page

Replicate it with one file

The whole process fits in the markdown file below. Hand it to an AI coding agent that can run commands, search the web and start sub-agents (Claude Code can do all three). Without sub-agents, the file has you run each brief in a fresh session. It interviews you first: your kids’ ages, your printer and its bed size, where you buy electronics, your budget, which AI services you have accounts for and what the kids already love playing with. Every later agent reads those answers. Then it runs the loop, with a brief template for each step.

A toy description with a price and a sample dialogue feels finished long before anything works, so the file ends with printing one toy and keeping a short play log until the third and fourth week, when the novelty has worn off.

handoff-toys.md

Download
# Handoff: run a toy workshop with a team of agents

You are an AI coding agent (Claude Code or similar) with shell access, web access and the ability to start sub-agents. Your user wants toys for their family that combine a 3D printer, cheap electronics and AI. Your task: run a design process with several agents that ends in a short list the family chose, a one-page presentation for each pick, and one toy printed and played with.

Everything you need is in this file. The process ran for one family in September 2026 and produced 53 toy ideas over two iterations. Its first round made mistakes (listed under "What goes wrong"), and the milestones below follow the repaired version. Work through them in order and verify each one before moving on.

## What you're building

```
interview ──► workshop/rules.md        (every agent reads it first)
                   │
research menu: what exists · fabrication · priced parts catalog
                   │
starter set  (the family's own ideas or prints, or one quick generator)
                   │
        ┌──────────┴──────────┐
     Remixer               Explorer
  remixes the set,       researches gap areas,
  never reads the        sees the set only at a
  research               final duplicate check
        └─── cross-breed ─────┘
                   │
         blind jury (5 criteria × 1–5)
                   │
   rework under "no re-equipping" ──► jury again
                   │
     family review: want / maybe / no
                   │
   one-pagers for the picks ──► print one, watch weeks 3–4
```

You are the lead: you write the briefs, start the sub-agents, check their files and make the cuts. In the original run the second iteration, from the two generators to the second jury, took about one morning of agent time. Apart from the agents' own usage, nothing costs money until the family decides to build.

## Ground rules (non-negotiable)

1. Children's data stays in the house. Names, voices and photos of a child go to a cloud service only if its terms allow minors and the parent agreed, and they never go into a public page or a repository with a remote. Files use age labels ("Kid A, 9").
2. You never buy, order, send, or start a print. You prepare; the user acts.
3. Every sub-agent reads `workshop/rules.md` first and writes its result to the file you name (`<role>-out.md`; some harnesses refuse sub-agent writes to files named `report*`, `summary*`, `findings*` or `analysis*`). Its chat reply is only a pointer to that file.
4. Nobody grades their own work. A fresh jury scores the generators' toys, and the family decides.
5. Verify every milestone against its files: count the toys, check every card has every field, open the pages.

If your tool can't start sub-agents, run each brief in a separate session with a fresh context. What matters is that each agent sees only the inputs its brief names.

## Milestone 0: Interview your user (~15 minutes)

Ask one question at a time and write the answers into `workshop/rules.md` (template in Appendix A).

1. Who plays, and how old are they? Ages are enough; keep names out of every file.
2. What do the kids pick up by themselves, what gathers dust, what do they do outside, and what are the adults into? Has anyone in the family designed or printed a toy? Ask for the files. In the original, the one toy the parent had designed and printed that year was a table-hockey game with no electronics, exactly the kind of toy the first round never proposed.
3. The printer: model, build volume, reliable materials, multi-material or not, the thinnest wall they trust, and whether they'll pause a print to drop in a magnet.
4. Electronics: soldering, boards already in a drawer, a computer at home that stays on (optional; it can host a local gateway).
5. The two or three shops they buy from, and a budget: a soft cap for a toy's minimal version, a hard cap for the full one.
6. Which AI services they have accounts for, and whether a home machine can run local speech and vision models.
7. Phone or desk for reviewing, the language of the cards, and whether the kids vote.

Optional but recommended: a week of observation before Milestone 3, one sentence per evening in `workshop/play-log.md` on what the kids played with and for how long. The original's independent process agent recommended it; the run skipped it.

**Provider terms for minors.** For every AI service a toy might use (language model, speech, image, text-to-3D), open its current terms and usage policy and add a row to `rules.md`: minimum age, whether a parent may operate it for a child, bans on content involving minors, region limits, URL, date checked. Terms change, so never fill this from memory. In September 2026 the original check found all of these patterns: a model API and a voice service that barred users under 18, a model that allowed minors with safeguards, a text-to-3D service that required 14+ and parental consent, an open 3D model whose license excluded the EU, and a text-to-speech service that banned output involving minors. Where the terms are unclear and the toy is for kids, use a local model or keep the AI in the workshop, run by the parent.

Verify: every template section is filled in; every provider row has a URL and a date.

## Milestone 1: Shared hard rules

Finish the hard rules in `rules.md` from the template, adjusted to the interview. Three of them come with their reason:

- At least two of three pillars (printing, electronics, AI), with AI allowed to work in the workshop: it generates a track, a level or a part's dimensions, and the toy itself can be purely mechanical. The original first round made all three mandatory, and 42 of its 43 toys needed electronics plus a running computer.
- Code holds the state. The model talks and picks from a whitelist, and its structured output is validated before it moves a servo, opens a lock or saves anything. One of the original ideators wrote the reason down: "a model can be talked around, code can't."
- Every toy has a minimal playable version, buildable in an evening or a weekend, and "a simpler version is an equal result". Without that rule, the original's second pass added features to almost every toy (Milestone 6).

The motif cap (each overused motif may be the core of at most two toys per generator) starts with the original's list: AI as narrator or voice, thermal printer, NFC tokens, the home gateway. Milestone 2 adds whatever your starter set overuses. Every card follows the format in Appendix B; its lines "The first 30 seconds" and "Why a tenth time" are the ones the jury scores first.

## Milestone 2: Research as a menu

Start three researchers in parallel. Their reports are optional material: nobody downstream has to use them.

- `world-out.md`: what exists where AI meets physical objects, interaction patterns, failures, 10 seeds. The original's key finding came from here: the strongest work uses the object as the model's input or output, and kids quickly see through a toy that only talks, then start insulting it.
- `fab-out.md`: print-in-place and compliant mechanisms, parts inserted during a paused print, tolerances, toy safety, text-to-3D and code-to-CAD, 10 seeds.
- `parts-out.md`: 50–80 priced parts from the user's shops, plus surprises. The original listed 68 items from two Czech shops, with surprises like a thermal receipt printer (928 CZK), a presence radar (198 CZK) and an 8×8 time-of-flight sensor (458 CZK).

```
You are a researcher for a family toy workshop. Read workshop/rules.md first.
Topic: <world | fabrication | parts>. Scope: <the bullet above, filled in>.
Use web search and fetch; for shop catalogs, your own headless browser if needed.
Output workshop/research/<topic>-out.md: a 5-sentence summary; findings with a URL
for each (parts: table name | price | in stock | link | checked on | toy use); what
failed or surprised you; 10 one-line seeds, each citing its finding. Say what you
could not load. Do not design finished toys. Write as you go.
```

**The starter set.** The Remixer needs a set to remix: the family's own list of ideas, the toys they've printed, or else one starter generator. One is enough, since a second one with the same inputs mostly repeats the first (see Milestone 3).

```
You are the starter generator. Read workshop/rules.md, then all of workshop/research/.
Write 15-20 toys in the card format: 10 quick first-hand ideas (bold is fine) and
5-10 grown from one seed each by a named operation (combine two patterns, another
audience, invert roles, one object into a swarm, change the sense, a toy that grows
with the child). A third must use AI abilities from the last two years. List the
building blocks your toys use (id, name, one sentence each) and name them on every
card. End with your top 5. Output workshop/starter/starter-out.md.
```

**Map the gaps.** Give this to a fresh sub-agent that wrote none of the starter set (in the original, the pattern was found by an agent that never saw the lead's notes). It counts the set and writes `workshop/gaps.md`. Per toy: does it need a computer or network, does the AI mainly talk, where is it played (table, house, outdoors), the play type (competition, chance, role play, physical sensation, making, puzzle), audience, minimal-version price. Thin cells become the Explorer's gap areas; motifs that appear everywhere go into the motif cap. The original set of 43: the AI mainly talked or wrote in 32, only 3 were played outdoors, the median price was 1,675 CZK (about €67), and music, movement, mechanical toys, puzzles and toys for teens or adults were nearly absent.

## Milestone 3: Two generators with different inputs

In the original first round, two ideators got different personas (a "free thinker" and a "grounded" engineer) and the same brief with the same three reports. They chose 4 of the same 5 seeds, gave 3 ideas identical names, and 40 ideas came down to about 34 distinct ones, all on one technical skeleton: a microcontroller, a gateway on the home computer, a cloud model, a voice. The lead's note afterwards: "independence was only in the persona." Design research calls this design fixation: people copy the examples they saw before they started, flaws included (Jansson & Smith, 1991).

| | Remixer | Explorer |
|---|---|---|
| Reads | rules.md, gaps.md, the starter set, the parts catalog, the family's own prints | rules.md and its brief (with the gap areas pasted in) |
| Never reads | the world and fabrication reports | all research, the starter set, the rest of `workshop/` until its duplicate check |
| Output | 15–18 toys `R01…`, 8+ building blocks | 15+ building blocks, 15–18 toys `E01…` |

Run both in parallel. In the original, 17 of 18 Remixer toys and all 18 Explorer toys play in their minimal form without the computer or network, and the Explorer's median minimal version cost about 425 CZK (about €17).

```
You are the Remixer for a family toy workshop. Read workshop/rules.md first.
Material: workshop/starter/<file> (read all of it); workshop/research/parts-out.md as
a menu; <the family's own prints> as an example of what they really print.
Do NOT read the world or fabrication reports: remix finished ideas instead of
starting again from the same material.

Six operations, each producing at least 2 toys:
1. Simplify to the core: keep the fun moment, remove the rest. Ideally no computer,
   no network, under <soft cap>, still two pillars.
2. Combine unexpected pairs: two toys, or a toy and a block from another category.
3. Transplant a strong block into a gap area from workshop/gaps.md.
4. Invert roles: the child teaches and the AI learns; the toy asks; the AI is the
   opponent; the output is a physical object instead of speech.
5. Improve without inflating: neither pricier nor more complex.
6. Split into a series of parts or levels that can be printed later.

15-18 toys (R01...), Origin line = source ids + operation; at least half play
without a computer or network; spread across <audiences>; 8+ new building blocks.
Output workshop/gen/remixer-r1-out.md: a 3-5 sentence intro on what you saw in the
set, toys, blocks, your top 5. Write as you go.
```

```
You are the Explorer for a family toy workshop. Read workshop/rules.md first.
You get different inputs from everyone else on purpose. Do NOT read workshop/research/,
the starter set, or anything else in workshop/. The only exception is step 4.

1. Quick research of your own, 10+ loaded sources, kept brisk (it's fuel):
   <gap areas from gaps.md>; <family hobbies>; printable kinetics (print-in-place,
   compliant and bistable mechanisms, automata, marble machines, spinning tops);
   music and rhythm; movement and skill; outdoors without Wi-Fi; puzzles and party
   games for teens and adults; AI other than a talking voice (vision on one snapshot,
   pose estimation, sound classification, a model writing OpenSCAD, small local
   models, AI as opponent, referee or level generator); what makers actually print.
   Note each source and what you take from it.
2. 15+ new building blocks, each an atom that can seed several toys.
3. 15-18 toys (E01...), Origin line = area + blocks, spread across <audiences>.
4. Only now: read the names and pitches of the existing set (<file>). Differentiate
   or drop any toy with the same core, and say which.
Output workshop/gen/explorer-r1-out.md: intro, sources, blocks, toys, duplicate
check, top 5. Write as you go.
```

## Milestone 4: Cross-breeding

Each generator reads the other's full output. List the overlap pairs you see before you send the brief (the original had five, among them two marble computers and two locks).

```
Round 2: read the other generator's output: <path>.
1. Overlaps: <R05 x E05, ...> and any others. For each: how they differ, whose core
   is stronger and why, whether a merge beats both. Be honest with yourself too: if
   the other version is better, say so.
2. 4-6 hybrids (XR01... / XE01...) that give something neither source toy has. All
   rules apply. A hybrid must be more than the sum of both toys' parts; justify any
   parts cost above the pricier source toy.
3. 2-3 fresh ideas the other output sparked (full card format; they count toward
   the 4-6 new items above, not on top of them).
Output workshop/gen/<role>-r2-out.md. Do not change your round-1 file.
```

## Milestone 5: Blind jury

A fresh agent that wrote none of the ideas scores all of them. Language models rate text they recognize as their own higher (Panickssery, Bowman & Feng, 2024) and tend to reward longer answers (Zheng et al., 2023), so the brief says both out loud. A jury on a different model family is better still; the original used one family throughout.

```
You are an independent juror for a family toy workshop. You invented none of these
ideas and the author does not matter. Judge the idea, not how well it is written.
A longer or more enthusiastic description must not get a bonus.
Read rules.md, the four generator files <paths>, and the starter set's names.

Score every toy 1-5 on:
1. The first 30 seconds: instantly clear what the player does, and tempting?
2. Why a tenth time: a real reason to return once the novelty fades (2-4 weeks).
3. Buildable at home: parts real and in stock, prices right, prints fit <volume>,
   minimal version really an evening or weekend. Web-check at least 5 prices or
   parts where something looks off, and list what you checked.
4. Novelty vs. the set: flag duplicates and near-duplicates.
5. Fun-to-complexity: penalize re-equipping; simplicity is a plus.
A hard-rule violation (pillars, AI allowed for the kid's age, safety) is STOP, with
the reason. Also ask who can make new content for each toy: if only the parent at the
computer can, the reason to return depends on the parent's free evening.

Output workshop/jury/jury-r1-out.md: table ID | name | 5 scores | sum /25 | STOP? |
one-line verdict; duplicate clusters with a winner; top 12 with reasons (no three
variants of one idea); one piece of advice per top-20 toy that must not add
complexity unless justified; what the pool is still missing.
```

The content question comes from the original jury's own finding. That jury also checked claims on the web: a 12 V pull magnet cost what the card said but drew 2 A, so four at once would overload the planned 12 V/2 A adapter. Expect the top 12 to lean toward the Remixer: 8 of the original's 12 came from it, because simplifications score well on buildability and fun-to-complexity. The Explorer's odder ideas still reach the family in Milestone 7.

## Milestone 6: Rework with "no re-equipping", then the jury again

In the original first round, the lead sent 16 finalists back to their authors for a second pass. Almost every prompt added something (masks, a motion wand, buzzers, a radar). Self-rated engagement went up, self-rated feasibility went down, and the total parts cost rose 32 %; one toy went from 2,000 to 5,940 CZK. Nothing got built.

```
Round 4: read workshop/jury/jury-r1-out.md in full. Its duplicate rulings are binding.
A. Rework your toys from the jury's top 20: <IDs>. Apply the advice and corrections,
   above all safety notes and wrong prices (<concrete corrections>). If you disagree,
   say why and keep the original. The rework must not increase the complexity or the
   price of the minimal version; if you can't avoid it, justify it in one sentence.
   Add "Changes after jury:" (1-3 sentences). Max ~25 lines per card (phone).
B. 2-3 new toys (NR01... / NE01...) for these gaps from the jury: <gaps>.
Output workshop/gen/<role>-r4-out.md. Do not change earlier files.
```

Then run the jury again on the changed and new cards, asking also whether the rework made anything worse. From here on, every card that changes goes back through the jury before anyone builds a page, including after the family redirects it. In the original, the jury after the parent's redirects caught some of the run's most concrete defects. A pocket puzzle's gravity lock would never lock: its 6 mm ball sat 0.7 mm into the drawer's notch, so pulling the drawer lifted it like a wedge with about 1 g of force. And a bridge toy's load test was meant to fill a 40 ml cup with marbles until the bridge broke, but the cup held about ten and filled first; the fix was a bag of rice, weighed afterwards.

Verify: every reworked card has its "Changes after jury" line, and no minimal-version price rose without a reason.

## Milestone 7: The family picks

Build a review surface for wherever the parent reads. The original was a small phone web app behind a login. A simpler version is enough: convert the final cards to `workshop/review/cards.json`, inline that data into one self-contained HTML file (no fetch, which browsers block for local files), and send the file to the parent's phone or serve it on the home network (`python3 -m http.server`). Keep decisions in localStorage, which lives on one device, and add an Export button. It needs a three-line summary per toy that expands to the full card; a verdict (want, maybe, no); checkboxes for liked building blocks and a note; for want and maybe, one next step (build a minimal version, a new toy on related blocks, three directions, other); a progress counter; and a Markdown export of want and maybe, which is the input of the next round.

Put the whole set in it, including every idea you cut. The original parent first picked 10 of 25 candidates from a list; two days later the verdicts on all 53 toys in the app replaced that pick, which had been "made without the context of the other toys". Half of the final 16 were first-round ideas the lead agent had not shortlisted.

If the kids vote, collect their votes before they see agent scores. When the family gives a toy a direction, send the card back to its generator with their words quoted verbatim, then to the jury. Toys built on existing models get a search of the big model-sharing sites: link, license, downloads, print results from the comments. One license in the original forbade rehosting, so the page only links to the model.

## Milestone 8: One-pagers for the final picks

Start page builders only on final cards: in the original, cards that changed during page building cost several repair rounds. One builder per 3–5 pages.

```
You build one-page presentations for toys from a family toy workshop. Read rules.md
first. Toys: <IDs and card paths>. Take content from the card faithfully; invent
nothing except the illustration and play situations that follow from the card.
Each page: a hero with name, pitch, age and the toy drawn as inline SVG (shape,
printed parts, where the electronics sit); an offline interactive demo of the core
loop (sound only after a click); the first 30 seconds and why a tenth time; what
printing, electronics and AI each do, with the code/model boundary visible; the
minimal version with a parts table, total and print time; safety and "watch out
for", including the jury's defects; the jury score out of 25.
Design: each page its own visual world from the toy's own world (a radio gets
bakelite and a tuning dial). Fonts must contain every glyph of <language>. Avoid
generic AI looks: cream + serif + terracotta, purple-blue gradients, emoji icons,
everything centered.
Contract: one self-contained HTML file; scripts only from <one pinned CDN>; no fetch,
no iframes; works at 360-400 px with 16 px side margins and no horizontal scroll;
touch targets 44 px; reduced motion respected; under ~300 KB.
Self-check in your own headless browser (e.g. Playwright): screenshots at 1280 and
390 px, scrollWidth <= innerWidth at 360, zero console errors. One fix round.
Output workshop/pages/<slug>.html and workshop/pages/<your-name>-out.md.
```

Then one more agent reads every page against `rules.md` and its card. In the original a page for ages 4 to 9 drew a cloud voice into its architecture diagram, against the family's local-voice rule for kids, and a reviewing agent caught it.

Optional public gallery: build it from a separate folder, and give its build a privacy gate, a script with test fixtures that fails on any family member's name, the home village, the surname, private links and e-mail addresses. Public data copies only verdict tiers and picked blocks, notes stay private, and kids appear by age ("my 9-year-old").

## Milestone 9: Print one

The original's process agent counted the first round: about 11 agent runs, one day, 0 CZK, zero hours of the kids' time, 43 ideas on paper, and the parent's print setup mentioned 0 times in the first-round files. It warned that a purely computational second round "would give nicer paper than the first iteration, but paper again", and it proposed a behavioral test: a toy works when a kid picks it up unprompted in week 3 or week 4. When this file was written, none of the workshop's toys had been printed yet.

1. Take the cheapest minimal version among the "want" picks. The original's cheapest finalists: a spinning-top arena at about 130 CZK (€5), a bridge kit at about 372 CZK (€15).
2. Print test pieces first for anything the jury marked unverified. The bridge kit's 0.15 mm skewer-joint clearance was untested, so a test comb with 3.0–3.4 mm holes comes first.
3. Help the user slice and assemble it. The user starts the printer, with an adult around while it runs.
4. Put it on the table without a pitch and keep the play log: who played, how long, did anyone ask for it.
5. Read the log in weeks 3–4. What the kids pick up by themselves earns a better version; what they ignore goes back to "maybe".

## What goes wrong

- Agents with the same inputs converge whatever persona you give them: 40 ideas, about 34 distinct, three identical names.
- Second passes add parts: +32 % parts cost across 16 finalists while feasibility fell. The "no re-equipping" rule and a jury after every change hold it back.
- Model judges favor longer, more enthusiastic text and text like their own. Keep authors off the jury and make it check prices on the web.
- A jury can compute that a lock won't lock; it can't tell you whether a 6-year-old reaches for the toy on day 20.
- The family's picks can differ a lot from the agents' ranking, so show them everything.
- Kids' data stays in the house: age labels in files, local models for children's voices, a privacy gate before anything goes public.

## Appendix A: rules.md template

```markdown
# Toy workshop rules (every agent reads this first)

## Audience (labels only, no names in any file)
| Label | Age | Loves | Ignores |
|-------|-----|-------|---------|
| Kid A | 9   | ...   | ...     |
Adults' hobbies: ...

## Hardware and budget
Printer: <model>, volume <X x Y x Z mm>, materials <PLA, PETG>, avoided <TPU>,
multi-material <yes/no>, min wall <1.2 mm>, pause for inserts <yes/no>.
Electronics: soldering <yes/no>, on hand <boards, sensors>, home machine always on
<yes/no>, local speech/vision possible <yes/no>.
Shops: <shop 1>, <shop 2>, <overseas, with delivery time>.
Budget: minimal version <= <soft cap>, full version <= <hard cap>.
Cards in <language>, reviewed on <phone/desk>, kids vote <yes/no>.

## Hard rules for every toy
1. Buildable at home: every part from the shops above or the printer above.
2. At least two of three pillars (3D printing, electronics, AI). AI counts when it
   works in the workshop and produces something the toy needs.
3. Kids and AI: only services allowed for the kid's age (table below); local speech
   and vision where possible; push-to-talk only; code holds the state, the model
   talks and picks from a whitelist, output validated before it moves, opens or
   saves anything; no images generated on a child's request; a basic offline mode.
4. Safety: no small parts for <youngest> (a part that fits the 31.7 mm small-parts
   test cylinder is small); magnets enclosed; no loose coin cells; lithium cells
   padded, protected and charged by an adult; heated or cooled surfaces capped in
   code and by a hardware cutoff that doesn't depend on the firmware; no mains
   voltage inside a toy; no unattended printing.
5. Every toy has a minimal playable version for one evening or weekend. A simpler
   version is an equal result. If it needs the computer or network, say what is
   left without them.
6. Motif cap, at most two toys' core per generator: <AI as narrator or voice>,
   <home gateway>, <overused motifs from gaps.md>.

## AI services and minors (checked by the agent, never from memory)
| Service | Used for | Min. age | Parent may operate? | Bans on minors' content | Region | URL | Checked |

## Building blocks
id (kebab-case) · name · category (printing | electronics | ai | mechanic |
interaction | architecture) · 1-2 sentences
```

## Appendix B: Card format

```
### <ID>: <Name>
Pitch: one sentence.
Who and when: age / who with whom / where / how long.
The first 30 seconds: what the player concretely does.
Why a tenth time: what brings them back next week.
Pillars: printing: ... · electronics: ... · AI: ... (at least two)
Minimal playable version: what to build, build time, price.
Full version: ... (optional)
Parts: main parts with prices, total.
Building blocks: existing ids + new ids.
Origin: source toys or blocks and the operation (or "new area: ...").
What's new: 1-2 sentences on why the set doesn't have this yet.
Presentation hook: how an interactive page would show it best.
```

## Appendix C: Folder layout

```
workshop/
  rules.md  gaps.md  play-log.md
  research/   world-out.md, fab-out.md, parts-out.md
  starter/    starter-out.md (or the family's own list)
  gen/        remixer-r1/r2/r4-out.md, explorer-r1/r2/r4-out.md
  jury/       jury-r1-out.md, jury-r2-out.md, ...
  review/     cards.json, the review page, the export
  pages/      one HTML file per final pick, builder notes
```

## Appendix D: Troubleshooting

| Symptom | Fix |
|---------|-----|
| The two generators return similar toys | Their inputs overlap. Check what each read; the Explorer sees nothing but its brief until the duplicate check. |
| Nearly every toy needs a computer and Wi-Fi | The pillar rule or the motif cap is missing from `rules.md`. |
| Reworked cards got longer and pricier | Resend with the "no re-equipping" rule quoted; compare minimal-version prices before and after. |
| A sub-agent says it's done and there's no file | It answered in chat. Give it the path again and ask it to write the file. |
| Accented letters in a display font look different | The font lacks those glyphs and the browser fell back. Check coverage for the card language. |

Nothing is printed yet

The agent that diagnosed round one had warned me that another round on the computer would give me “nicer paper than the first iteration, but paper again”, and it was right. The runes, the stone and a spinning top arena are the three toys getting off the paper.

The other 13 are on toys.malacka.cz, among them a family radio you tune to a year, a barrel organ that tells a fairy tale as you crank it and prints a picture at every fork in the story, and a guardian you can try to talk out of his password.