# Using imagegen to design working products

Practical field guide for the next agent. Written 14 September 2026 from the
Gyro, Eiko and Meryl design sessions, screenshot-repair experiments and onboarding
implementation in this repository.

“Imagen” was used conversationally for our image-generation tool. This guide
does not establish a particular vendor model, model version or diffusion architecture.
It records how we used the available built-in imagegen tool with vision, code,
browser inspection, user feedback and independent engineering review.

**The central lesson: generate possibilities, implement deliberately, and test
the actual product. A beautiful image is neither a specification nor a verdict.**

Follow the current user scope, tool instructions, permissions and delegation
rules. This guide is experience to apply, not authority to ship a concept,
spawn reviewers, change another site or run a transaction.

## 1. Start here: the short workflow

1. Read the product intent, current implementation and approved visual references.
2. Pick one user task and one concrete friction point.
3. Capture the relevant real screen and record its state.
4. Generate alternatives with distinct hypotheses, or request one focused repair.
5. Inspect each image. Record what to adopt, adapt, reject and leave unchanged.
6. Implement a small coherent slice in the appropriate HTML/CSS/component layer.
7. Verify meaning, behaviour, accessibility and actual responsive geometry.
8. Capture the implemented result. Use another imagegen pass only if it answers
   a remaining visual question.
9. Stop when there is no specific useful change left. Publish the comparison,
   prompts, decisions and test evidence for review.

For an established screen, begin at the component level. Do not default to three
whole-app redesigns just because that worked during early exploration.

For a new product direction, a wider exploration can be appropriate. The user
must still choose or authorize the direction before a concept becomes a release.

## 2. Divide the responsibilities

| Participant/tool | Useful contribution | What it does not establish |
|---|---|---|
| Image model | Composition, visual hierarchy, brand directions, illustrative alternatives | Correct numbers, actual controls, responsive behaviour or product understanding |
| Agent using vision | Compare screens; spot omissions, grouping problems, drift and contradictions | Exact CSS geometry or what happens after a click |
| Agent using code and browser tools | Implement; measure; inspect state; test real handlers and failure paths | Whether an unfamiliar human understands the product |
| Fresh engineering reviewer | Challenge assumptions and check implementation independently | Independent user research, merely by being a fresh model context |
| Product owner / representative users | Approve intent, judge usefulness and expose comprehension problems | Automatic correctness of every implementation detail |

Keep these responsibilities separate. Asking an image model to redraw a screen
and then treating its own polished annotations as approval is circular.

## 3. What happened across these sessions

These are observations from this project, not universal benchmark results.

| Session | Useful result | Failure or limit | Lesson |
|---|---|---|---|
| Eiko A2/B2/C2 exploration | Existing collateral and brand guidance produced directions the user preferred; alpha combined workspace, position emphasis and optional dark styling | A visually approved reference did not mean every label or interaction was approved | Start from approved identity and intent, not from an unbranded blank canvas |
| Eiko phone records repair, 11 Sep | Suggested insets, padded headers, a 2×2 fee group and LP separation | Added a duplicate disclosure arrow; resynthesized fonts and geometry | A focused crop can be productive; native disclosure behaviour still needs code review |
| Eiko desktop strikes, 11 Sep | Suggested removing the enclosing border and separating choices | Still drew unequal widths and kept the offset label | Extract the relationship, then implement equal grid columns and measure them |
| Gyro user-first layout, 12 Sep | Larger actions and personal position before global statistics | Generated mobile scale was not literal; Poke stayed too prominent and explanations repeated | Prioritize the user's action, not the protocol's headline metrics |
| Meryl design competition and refinements, 12 Sep | User chose A's continuous white/cobalt statement; later fixes aligned fees and simplified context | Missing/duplicate strikes, tiny targets, extra menu affordances, stubborn punctuation and misleading exit terminology | Preserve the coherent direction, but verify every detail; do not keep editing a raster to fix what belongs in code |
| Three-site cycle, 13 Sep | Nine concepts plus three screenshot-diff proposals reinforced clearer holdings, less nested framing and shorter seller guidance | Invented history, omitted controls, tiny phone UI and repeated suggestions; later gains were modest | Breadth helps exploration; convergence needs stricter scope and a stopping rule |
| Onboarding study, 14 Sep | Intent grouping plus an optional selected-path explanation clarified two destinations and three actions | A map invented a reverse liquidity flow; phone variants dropped buyer/seller prerequisites | Diagrams carry claims. A wrong arrow can teach the wrong product more persuasively than text |
| Gyro onboarding implementation, 14 Sep | Native diagrams, optional steps, skip/reopen behaviour and a browser-local returning-user preference | A literal stacked implementation was long on phones; it needed a compact native layout | Generated desktop/phone boards do not remove the need to design real responsive behaviour |

The record supports “this process found useful changes.” It does **not** yet
support a measured claim that onboarding improved task completion or conversion.
The proposed newcomer comprehension test has not been performed with users.

### An important mistake came from our prompt

The original Meryl competition prompt supplied the label “Estimated exit.”
Implementation review established that the holdings field could be a spot mark,
not a verified executable single-asset exit. We retained the correct Holdings
meaning instead of copying the prompt or image literally.

Do not blame every semantic error on imagegen. If the agent supplies a misleading
brief, the model can reproduce it faithfully. Validate the words and data you
give the model before evaluating how well it followed them.

## 4. Define a question before requesting an image

A useful brief contains:

- The user and their immediate task.
- The observed obstacle, with a screenshot or concrete example.
- The product meaning that must survive.
- The permitted scope of change.
- What would count as an improvement.

Weak: “Make this app look more professional.”

Better: “On a phone, the seller must scroll through a long explanation before
reaching IV and quantity. Keep the existing form and gates; make its next action
easier to find and move supporting mechanics into optional detail.”

Better: “A newcomer confuses Vault with Market. Can they identify where to
deposit assets, buy an option, or list existing vault shares without coaching?”

A useful acceptance criterion might be “all role choices are discoverable and
the seller prerequisite survives on phone.” “Looks cleaner” is not enough.

### Give variants different hypotheses

For the onboarding study, the alternatives were:

- **Intent choices:** route people by what they want to do.
- **Product map:** explain the relationship between Vault and Market.
- **Optional guided preview:** teach one choice after the user selects a path.

Those are meaningfully different approaches. Three palettes applied to the same
layout would not test that question.

Do not average the most attractive pieces of every concept into one screen.
Choose a coherent layout and borrow only improvements compatible with it.

## 5. Prepare references and honest captures

### Use approved collateral without making it a cage

Supply the site's actual logo, visual guide, relevant illustrations and approved
screens. Identify their roles explicitly:

- Image 1: actual screen to edit.
- Image 2: brand reference only.
- Image 3: another state for context, not another screen to reproduce.

In Eiko, the approved collateral mattered. In Meryl, the reference image carried
asset names and identity, while the new direction intentionally changed the old
Gyro-derived palette. A reference can constrain identity without freezing layout.

### Record the capture state

At minimum record:

- Site, route and revision/resource versions.
- Viewport in CSS pixels, device emulation and scroll position.
- Chain, asset and strike, if relevant.
- Disconnected, approved read-only account, or clearly labelled fixture.
- Loading, ready, empty, error, paused and disclosure states.
- Any browser-only stabilization used for the capture.

Desktop and phone are two implementations to inspect, not merely two exports.
In our sessions, changing mobile emulation reloaded a page and removed injected
read-only account state. Some first phone captures were therefore disconnected
while the desktop had a position. We retained the distinction and made additional
populated captures; we did not silently call them identical-state comparisons.

For visual comparison, a frozen DOM fixture can help keep data steady. Label it
as a fixture, remove wallet/RPC behaviour, and do not use it as evidence that the
live app works. Separately test the real loading and recovery paths.

Browser-only pausing of a known refresh interval helped avoid capturing a
transient “Loading…” state after an actual read. This was not a production polling
change. Do not clear every timer or fabricate ready values to obtain a pretty image.

Prefer an approved public account and a real read-only mode for populated views.
Never connect or sign merely to stage a screenshot. Do not label synthetic
wallet balances as real holdings.

## 6. Prompt for the right kind of output

### A. Exploration recipe

```text
Use case: ui-mockup.
User/task: [who is trying to do what].
Observed problem: [specific friction].
Inputs: Image 1 is [role]; Image 2 is [role].
Hypothesis: [what this variant changes and why].
Keep: [brand identity, product meanings, essential controls and states].
May change: [layout/hierarchy/illustration boundaries].
Output: [one component / desktop and phone concept board].
Avoid: [the few relevant failure modes, not an unrelated list].
No invented product capabilities or financial data.
```

### B. Focused screenshot-repair recipe

```text
The supplied screenshot is the edit target.
Change only [component]. Improve [alignment/grouping/scan order].
Preserve [exact values, signs, units, selected options and visible controls].
Preserve [expanded/collapsed, disabled, empty or paused states].
Keep surrounding content unchanged.
Return one revised screen, without a device frame or comparison annotations.
```

### C. Post-implementation review recipe

```text
Image 1 is before; Image 2 is the actual implementation.
Compare them for [user task]. Different live values are not design changes.
Suggest at most three concrete, buildable improvements, if any.
Keep [critical meanings, controls and states].
Do not invent a problem simply to produce another variant.
If redrawing, treat the implementation as the edit target.
```

These are starting points, not incantations. Read the current tool/skill
instructions before use. The full historical prompts are linked at the end.

### What to constrain tightly

Constrain facts: product rights, units, signs, prerequisites, required controls,
selected states, conditional transitions and what is actually available.

Allow exploration in composition, hierarchy and grouping. If every old layout
decision is frozen, the model cannot propose a meaningful alternative.

Some historical prompts were very long, dense with exact figures and prohibitions.
That helped enumerate invariants but did not guarantee compliance. A better next
experiment is a smaller target, clearer priorities and a short content contract.
That recommendation is not a measured prompt-performance result.

## 7. Inspect outputs for meaning before beauty

For each output, keep a short decision table:

| Observation | Decision | Reason / check |
|---|---|---|
| Position becomes the leading element | Adopt | Matches the holder's first question |
| Supporting content uses a simpler grid | Adapt | Implement with actual responsive sizing |
| New chart or historic row appears | Reject | No underlying data or feature supplies it |
| Existing controls disappear | Reject | Less clutter achieved by lost functionality |
| Current layout already solves the problem | Keep current | No concrete benefit from another edit |

Audit these recurring failures:

- **Numbers:** invented epochs, changed signs, rounding, token-case changes,
  duplicated or omitted strikes, invented totals and positive-looking results.
- **Meaning:** holdings relabelled as guaranteed exit; fees presented as profit;
  listing IV confused with buyer price; prerequisite omitted.
- **Controls:** refresh removed, duplicate chevrons, new dropdown arrows with no
  action, invented menus, changed disclosure state.
- **Layout:** phone UI is a scaled desktop, labels shrink, cards repeat, whitespace
  grows without improving the task, a control looks disabled when available.
- **Claims through imagery:** a permanent stream of incoming coins implies
  dependable income; a circular arrow can imply a guaranteed lifecycle.

Onboarding B drew “Sold shares become liquidity” as a reverse flow. We rejected
the map as drawn. A initially said “Each action is independent,” conflicting with
the seller's need for vault shares. Its phone version dropped the buyer's
no-deposit cue. These were semantic failures, not small styling defects.

After a targeted revision, both device views retained the prerequisites. The
model still kept an ETH icon despite a request for abstract asset art. We recorded
that remaining issue and used generic native diagrams in code; another raster
edit was not necessary to make the implementation correct.

## 8. Implement relationships, not raster pixels

Translate the useful design idea into a layout rule:

- “These choices should feel equal” → equal-width responsive grid tracks.
- “This is one account statement” → shared grouping, aligned labels and values,
  fewer nested backgrounds.
- “This number leads” → intentional typography, not a new calculation.
- “This explanation is secondary” → a real disclosure with keyboard behaviour.
- “These routes have different requirements” → visible concise prerequisite text
  beside each route, including on phone.

Use HTML/CSS/SVG for controls, labels, diagrams with meaningful text, charts from
data, and existing icon systems. Raster art can help with mood, mascots or a
genuinely useful illustration. It should not become a screenshot-shaped UI whose
buttons, text, layout and accessibility no longer work.

We considered new generated UI assets during the three-site pass. Existing
artwork was sufficient. We did not add decorative images just to use imagegen.

### Keep the engineering boundary explicit

Site design must not silently replace financial logic or destroy multichain
capability. Reuse semantic hooks and existing actions; do not clone an entire
engine merely to adopt a skin.

Earlier user feedback was not “make Gyro look exactly like Eiko.” It was that an
approved, working reference had lost its visual treatment and engineering value
during porting; the user also reported an unstyled, misplaced share card. Preserve
the reference's useful capabilities without forcing the same brand. Investigate
stylesheet loading, resource versions, component boundaries and modal behaviour
before answering a broken implementation with another redesign. A reported cache
concern is a diagnostic lead, not a proven root cause.

Eiko alpha exposed real restrictions in the base: JS forced visible layout,
and LP status assumed a particular DOM wrapper/order. We added selective
presentation hooks and CSS variables. We did not rewrite the financial engine
for the design. Later, alpha was promoted while beta remained a frozen reference.

When fixing a UI state defect, preserve the underlying authority. The Sep 13
review found paused listing still looked enabled. The fix mirrored the existing
pause state onto a submit-only fieldset; it did not unpause a market or clear a
button's transaction-in-progress state. Cancellation, claims and withdrawal
remained outside that new fieldset.

### Separate three kinds of diffs

1. **Semantic diff:** did meaning, data, prerequisites or available actions change?
2. **Layout diff:** did order, spacing, alignment and visual emphasis improve?
3. **Behaviour diff:** do the real controls, states and responsive views work?

Raw pixel subtraction between a screenshot and a generated image largely
measures resynthesis of fonts, antialiasing and geometry. It is not a correctness
test. Pixel comparisons are more useful for controlled browser-to-browser
regressions, with dynamic content accounted for.

## 9. Keep words short without removing the product

The user's direction is concise, plain language—not a growing stack of caveats.
Do not replace investigation with a paragraph explaining why a wrong-looking
number might be acceptable. Fix the number or its meaning.

Keep information needed to decide: amount, unit, relevant comparison, prerequisites
and consequences. Put mechanics near the point where someone needs them, or in
optional detail. Do not hide essential price information behind a tooltip.

Useful distinctions from this project:

- Holdings are not automatically a verified withdrawal quote.
- Fees are one component of a return, not synonymous with profit.
- The approved return comparison here is against a passive limit order at the
  chosen strike. Do not replace that benchmark for aesthetic reasons.
- Buyers assess the total executable price; seller IV and buyer implied IV are
  not interchangeable. See the product intent for the actual ruling.
- A buyer does not need a vault deposit first; a seller needs vault shares.
- Listing an offer is not the same as somebody filling it.

These are repository-specific semantics to preserve, not a universal definition
of every product named Vault or Market.

## 10. Test the first visit and the established flow separately

The onboarding design was only half the task. The implemented flow also needed:

- An optional first-visit entry, not a forced tutorial over the user's task.
- Skip and an always-available way to reopen the guide.
- A browser-local preference, not wallet identity tracking.
- Direct-link and read-only-demo behaviour that preserves the user's purpose.
- Storage-denied and corrupt-preference fallbacks.
- Correct focus after opening, skipping and choosing a destination.
- Routing to the actual existing tabs without opening a wallet or submitting.

Gyro snapshots prior-use hints before app boot can write today's cache. Otherwise,
a new visitor could be misclassified as returning during the same page load.
The guide records skip or destination selection, not proof of understanding.
The explicit `intro=1` preview flag overrides normal suppression; test returning
behaviour on the plain URL, not on a forced-preview URL.

The first native phone layout stacked the illustrations too tall. We moved the
miniature diagrams beside their respective choices while keeping prerequisites
readable. This came from inspecting the browser, not from trusting the phone frame
in the generated board.

Thirty onboarding cases and the existing regression suites passed. That verifies
specific behaviour; it is still not a newcomer comprehension result.

## 11. Verification gates before handoff

### Product meaning

- [ ] Every financial label has an identified source and correct scope.
- [ ] No generated number, claim, arrow or prerequisite was copied unchecked.
- [ ] Negative results, genuine empty states and actual history remain visible.
- [ ] Buyer/seller roles and units are not conflated.

### Implementation

- [ ] Existing IDs, handlers, inputs, hidden states and entry gates survive.
- [ ] Check loading → ready, error → retry, pause → resume and context switches.
- [ ] New styles do not override disabled text, native disclosure markers,
  modal centring, focus indication or a busy transaction button.
- [ ] Measure actual responsive geometry, not raster dimensions. We used 44px
  minimum controls and 54px primary actions as project design targets.
- [ ] Test a narrow phone plus desktop and a relevant intermediate breakpoint;
  include long labels, small prices, selected/OFF states and populated positions.
- [ ] Inspect the actual loaded resource versions and computed styles.
- [ ] Keep scope local: unrelated forks, configuration, indexers and beta remain
  unchanged unless the user authorized those changes.

### Evidence

- [ ] Fresh captures reflect the implementation being handed off.
- [ ] Distinguish live read-only captures, static fixtures and generated proposals.
- [ ] Record what tests establish, what was not tested, and any known unrelated failure.
- [ ] No unrequested signing, transaction or public share publication occurred.
- [ ] Review gallery links and assets resolve on the user's actual host.

Do not turn this checklist into self-certification. A test based on the wrong
expected value can pass. An independent reviewer needs the source of the
expectation, not just a green result.

## 12. Tool and capture lessons worth carrying forward

These are session observations; check the currently available tool instructions
rather than assuming the same API or limits forever.

- The built-in tool returned an image and an output hint containing its saved
  path. Copy project artifacts into the repository; do not leave gallery assets
  pointing at an agent-private generated-images directory.
- Inspect local reference images before passing them to the image tool. A path
  alone is not evidence that the agent knows what it contains.
- One attempted request with six reference paths was rejected: this session's
  built-in tool accepted at most five. Reduce redundant context or split the
  task; do not quietly omit a required edit target.
- **Do not log the whole result object.** We did this in a parallel batch and
  dumped huge base64 strings into tool output, causing severe truncation and
  context waste. Render through the image helper; log concise path/status metadata.
- Parallel generations can reduce waiting, but extra variants are not extra
  independent evidence. Keep requests bounded, save their prompts and associate
  each result with its own input set before starting the next stage.
- Browser page IDs changed after reconnects, and separate tool contexts could
  expose different IDs. List pages in your own context and use dedicated review
  pages; never manipulate another agent's or the user's wallet session by guess.
- Selecting/focusing the target page improved screenshot reliability. Capture
  to an allowed path, then copy into the project when needed.
- Emulation can reload the app. Recheck the account, tab and readiness state
  after changing it; do not assume a loaded position survived.
- We observed stale-looking styling after editing/reloading. A resource-version
  bump made the new CSS visible. Check the served HTML, asset URL and computed
  styles before applying another layout patch to compensate for stale resources.
- The tool-result metadata inspected in these runs did not identify the backend
  model. An API model catalogue is not evidence of which model a particular
  built-in call used. Report only what the actual result establishes.

## 13. Stop rules and the next improvement to the process

Stop a visual loop when:

- The scoped friction is resolved and implementation checks pass.
- New proposals repeat changes already implemented.
- Differences are taste rather than task improvements.
- The model fixes one small detail while regressing required content elsewhere.
- The next useful question requires a human user, not another image.

An unchanged screen can be the right outcome. Meryl's established statement
needed much less work than Eiko's layout during the shared cycle.

For the next cycle, we recommend a blind before/after review where possible:
give the reviewer the task and invariants without praising the newer design;
allow “neither is better.” This is a proposed improvement, not a method proven
by the sessions documented here.

The highest-value missing evidence is a short observed newcomer test. Ask people
where they would go to deposit, buy without shares, and list existing shares.
Record wrong choices, hesitation and needed explanations. Preference for a mockup
alone does not establish that its product model is understood.

## 14. Save enough for another agent to continue

A useful dated artifact set is:

```text
site/img/designs/YYYY-MM-DD/
  before-desktop.png
  before-mobile.png
  concept-a.png
  concept-b.png
  pass1-desktop.png
  pass1-mobile.png
  designer-diff.png
  final-desktop.png
  final-mobile.png
site/DESIGN-ITERATION-YYYYMMDD.md
site/designs.html#dated-section
```

Use only the files needed for the actual experiment. Add buyer/seller, returning
or error-state captures when those are the task. Keep old variants distinct and
preserve archived gallery sections, including their requested hidden state.

The notes should contain the question, input roles, full prompts, tool mode,
capture conditions, adopt/adapt/reject decisions, implementation scope, checks,
remaining uncertainty and exact next question. Label a rejected concept where
it is displayed, not just in an obscure footnote.

When working over SSH, publish a usable gallery on the intended site and check
its links. Local images alone are a poor handoff when the user cannot open them.
Preserve unrelated working-tree changes and staged files; use scoped commits.

### Copyable next-agent brief

```text
Read IMAGEGEN-DESIGN-GUIDE.md, the product intent, and the relevant dated study.
Task: improve [one user journey/component] on [site only].
Current friction: [observable problem].
Approved direction and invariants: [specific references].
Allowed change: [presentation / explicitly scoped behaviour].
Capture [states/viewports] with [real read-only data or labelled fixture].
Use imagegen only to answer the remaining visual question.
Produce an adopt/adapt/reject decision, implement one coherent slice, and verify
the actual browser and required state transitions. Do not publish or transact
outside the user's authority. Save the comparison, prompts and checks.
Stop if another image would not resolve a concrete remaining problem.
```

## 15. Source trail

Read the relevant source rather than treating this synthesis as current runtime
configuration or a replacement for the user's latest ruling.

- [Product intent and buyer/seller ruling](.claude/rules/wheel-product-intent.md)
- [Eiko approved directions and archived first batch](eiko/designs.html)
- [Eiko alpha boundaries and later promotion](eiko/ALPHA.md)
- [Phone screenshot-repair experiment and exact prompt](eiko/img/designs/visual-fix/experiment.md)
- [Desktop strike repair and measured outcome](eiko/img/designs/desktop-fix/experiment.md)
- [Gyro 12 Sep triage, user override and implementation](gyro/DESIGN-TRIAGE-20260912.md)
- [Meryl competition, prompts and two refinement passes](meryl/DESIGN-COMPETITION-20260912.md)
- [Gyro 13 Sep cycle](gyro/DESIGN-ITERATION-20260913.md)
- [Eiko 13 Sep cycle](eiko/DESIGN-ITERATION-20260913.md)
- [Meryl 13 Sep cycle](meryl/DESIGN-ITERATION-20260913.md)
- [14 Sep onboarding hypotheses, rejected map and refined concept](gyro/ONBOARDING-STUDY-20260914.md)
- [Implemented first/returning onboarding behaviour](gyro/ONBOARDING.md)
- [Onboarding regression tests](gyro/test-onboarding.mjs)
- [Three-site design boundary and artifact tests](gyro/test-design-iteration.mjs)

Relevant checkpoints: `365488ea` (three-site design cycle and Eiko promotion),
`58a4c038` (onboarding study), `a77d0acc` (Gyro onboarding implementation).
These identify historical work, not a promise that the current head is unchanged.
