Deadwax
How we built it / Field notes 01Seattle, WA · An ongoing project

Five humans.
Eleven agents.
One record.

A question about which pressing to buy became an app for collectors, on the web and in both app stores. It also became a place to learn how people and AI can build, run, and grow something together.

Hope’s hand-drawn Deadwax skull mark

Hope drew the skull.
We built a project around the things we love.

5 Humans Taste, domain expertise, and every decision that matters.
11 Agent roles Defined scope, named owner, reviewable output.
3 Surfaces Web, iOS and Android, on one shared API.
6,800+ Pressing panels Panels that gave a collector usable pressing information, since May.
Jason, Hope, Caden and Chris outside Easy Street Records in West Seattle
Jason, Hope, Caden, and Chris at Easy Street Records in West Seattle. May 2026. Behind us is the Mother Love Bone mural, originally painted by Pearl Jam’s Jeff Ament.
Records. Software. Everything we learned in between.Start the story

We wanted to find
the right pressing.

Two copies of the same album can look almost identical and sound very different. The clues are there: the label, the mastering credit, the tiny letters scratched into the deadwax. Finding them takes work.

Discogs gives collectors an extraordinary database. The research around it is scattered across release pages, forum threads, and shootout videos. We wanted to bring that context together, alongside the records we already own and the ones we’re looking for.

That is the product. For me, there is a second reason to keep building it. Deadwax is where I test the engineering practices behind AI-assisted work: how to give an agent useful context, how to review its decisions, and how to know whether a change made the product better. Someone is standing in a record store, holding a copy of an album, waiting for an answer. That is a better test than a demo.

My daughter Hope drew our skull. Seeing that drawing become the app’s identity, and then a sticker we can hand to someone, is one of my favorite parts of the project. The code matters. So does making something with your family and friends.

We wrote down what kind of company we wanted to be, and those words go into the agents’ working context too. “Customers first” means checking whether a feature helps someone in a record store. “Data integrity” means saying what we know and where the evidence runs out. “Ship to learn” means listening when a collector tells us a feature made things worse. The principles become useful when they change a decision.

Hope’s original gold Deadwax wordmark, with a skull and record
Hope’s artwork became our visual identity.
A circular purple and black Deadwax skull sticker design
One of our sticker designs. Something to take beyond the screen.
The humansFive · judgment
Jason Founder & product lead

Product direction, priorities, the agent workflow, and the decisions about what we release.

Chris Pressing expertise

Collector judgment. The person who can challenge a recommendation that sounds convincing but misses the record.

Hope Visual identity

The brand, including the original hand-drawn skull mark. Brand changes come back to her.

Caden Community

Instagram, community voice, and helping the product find the collectors it is for.

David AI workflow advisor

Advice on agent workflows, context, and getting more useful work from the tools.

Humans own judgment. AI owns execution within the task we give it. That is the principle we work toward, with tests, reviews, and named human decisions around the output.

The work ends up
in someone’s record bag.

Collection and wantlist browsing. Pressing research with sources. Scanning. Stories that give you another reason to put a record on.

Deadwax Discover screen showing Seattle Sound and Psychedelic Underground in Collectors Corner
01 / Follow a sceneCollectors Corner connects records through places, labels, and musical movements.
Pressing Intelligence for Jar of Flies showing competing source views and two different editions
02 / Understand the copyPressing Intelligence presents recommendations, uncertainty, and the sources behind them.
Album Story in the Deadwax iOS app, with record history and source links
03 / Stay for the storyAlbum Stories adds the people, sessions, and accidents behind the music.

Actual iOS screens from September release work. Collectors Corner is live on web; the latest native update is in release preparation.

Take a picture.
Then read the tiny letters.

The barcode is easy. Older records often ask more of you.

We built ways to scan a barcode, read a label, and recognize an album cover. On iOS, you can also speak the faint deadwax markings into the microphone instead of typing them while holding the record. The app gives you a chance to check the text before searching.

That last idea came from my own record-store habit: turn the record toward the light and read the etching aloud. When Chris was surprised by the microphone turning on after a scan, we changed the flow to ask before enabling it. A small conversation became a product rule.

The camera can help identify the album. The label and the runout can help distinguish the pressing. Keeping those steps connected, without pretending the sleeve proves the exact copy, is the interesting design problem.

01

The sleeve

Recognize the album from its cover.

02

The label

Read the label and catalog details.

03

The deadwax

Speak the markings, check the text, find the pressing.

These are distinct recognition paths. The current cover-recognition implementation is on iOS; native barcode and label workflows also serve Android.

Shipping the app
changed the experiment.

The first challenge was getting something useful into people’s hands. The next was learning how to keep it useful.

We began with a fairly elaborate cast of agent roles: product, design, engineering, architecture, testing, marketing, and legal research. It helped us divide unfamiliar work. But the org chart was never the result. After launch, we spent more time simplifying the operating model and following the problems collectors actually reported.

  1. The web app gives us a starting point.

    Deadwax goes live on the web. The core idea is narrow: connect an existing Discogs account, browse vinyl, and make pressing research easier.

  2. The web app becomes a foundation for iOS.

    We choose to build a native iOS app against the same backend, carrying the collection and pressing-research work into a new interface.

  3. iOS reaches the App Store.

    The May 7 launch adds a new kind of work: store releases, device behavior, large collections, and feedback from people outside the project.

  4. We change how we run the work.

    The daily launch routine stops fitting a live product. We separate ongoing fixes from larger feature projects. Claude becomes the primary developer and orchestrator; Codex becomes a focused review and specialist layer.

  5. Android v1 goes public. We also undo a feature.

    Deadwax reaches Google Play. On iOS, we roll back batch scanning after collectors tell us the camera flow hides the pressing results they came to see. A feature can function correctly and still be the wrong experience.

  6. More reasons to explore. Better questions about AI.

    Collectors Corner grows to 18 curated scenes, with Album Stories and pressing context. The next engineering priority is an evaluation lab: compare models and workflows on the same real tasks, then decide what earns its cost.

Public milestones: App Store version history ↗ · Google Play ↗ · September Android launch post ↗

Writing code became
one part of the job.

I started by giving the agents jobs. Then I had to teach them what good work looked like here.

Eleven roles came out of that. Not eleven employees — eleven jobs with a defined scope, a named owner, and output somebody has to be able to review. People still decide what to publish, what to spend, and which legal or product questions need more help.

At first, much of this meant separate Claude and Codex sessions with instructions I carried between them. We moved more coordination into shared files and reusable workflows as the project grew. Today, Claude does much of the product development, Codex supplies independent review and selected specialist work, and Gemini helps extract pressing evidence from videos. The scheduled jobs have their own configuration. The team has changed shape more than once.

The role is only the starting point. Each agent needs the mission, the relevant product decisions, the design rules, the code it will touch, and the test that will tell us whether it has finished. That is context engineering. The surrounding tools, permissions, scripts, checks, and handoffs are the harness. Building that working environment has become a large part of the project.

I have used models such as Fable, Astra, Sonnet, and Sol for different kinds of work. We give subagents bounded tasks and route simpler work to faster, less expensive models, with stronger models for difficult reasoning and review. Shared context keeps each handoff from becoming another round of rediscovery. We record those choices in CLAUDE.md and AGENTS.md, then revisit them when the results or the cost stop making sense.

The agent rolesEleven · execution
Director / Orchestrator Claude

Task packets, dependencies, merge order, and the handoffs between the rest.

Product Manager Claude

Scope, priorities, user stories, acceptance criteria and the metrics we judge them by.

Designer & Researcher Claude

Collector workflows, visual rules, accessibility, and using a phone with a record in your other hand.

Architect Codex

Independent review. Reads the diff before a merge and blocks on real findings.

Backend Developer Claude

Authentication, API behavior, data pipelines and reliability.

Frontend Developer Claude

The web app, responsive workflows and interaction details.

Tester / QA Claude

Regression tests, browser automation, CI and the evidence behind a release.

Swift iOS Developer Codex

SwiftUI, Apple conventions and how the app behaves on an actual iPhone.

Android Developer Codex

Kotlin, Compose, Material 3, and native Android behavior.

Marketing Manager Claude

Positioning, go-to-market, community and launch material. A person still decides what we say.

Legal & Compliance Advisor Claude

Research on naming, privacy, licensing and API-use questions. Advisory, never the final word.

These are reusable responsibilities, not eleven continuously running employees. Claude does most of the product development and orchestration; Codex supplies the independent review and the two native specialties. Gemini’s video extraction and the claim judge are product-data and evaluation jobs, outside this roster.

Every change walks this line

Can stop the line

01

Worktree

The agent gets its own copy of the repository, cut fresh from main — the code and the written strategy it works from.

02

Local checks

The regression bundle runs on this machine first. CI minutes and deploys are money.

03Gate

Independent review

A second model reads the diff and blocks on real findings. Expect several rounds.

04

Pull request

Rebased onto main. Protected files are stripped before anything is pushed.

05Gate

CI

Security rules, lint, types, unit tests, browser tests, and the native builds.

06

Merge

Squash onto main, delete the branch. The agent never merges its own work unreviewed.

07

Deploy & verify

Backend, then the site and CDN. Then we exercise the changed route in production.

08Human

The human gate

A store release waits for a person. Chris and I use the build with a record in hand.

Some of these are machinery. Some are discipline. A script resets our strategy documents out of a change before it can become a pull request, and a check fails the pull request if they survived. Another refuses to tell a collector their bug is fixed until a person has confirmed it. The rest — merging only on green, waiting for a human before a store release — are habits we hold ourselves to. Knowing which is which is the part that matters.

Write the script first.

When work has a repeatable answer, we write a script. We add tests for the script, run them whenever it changes, and have the agent use that tested script. Review preparation, test runs, branch cleanup, deployment, and pressing-data validation all follow this pattern. It saves the agent from rebuilding the same sequence of commands and gives us something we can inspect and improve.

That rule lives in CLAUDE.md and AGENTS.md: deterministic work belongs in scripts. Skills guide the parts that need judgment. Before an agent starts improvising, it should check whether we already have a tool for the job.

Make the checks inspectable.

Our everyday evaluation of code quality starts with tests. Unit tests check the logic. Playwright walks through browser workflows. GitHub Actions runs those checks alongside lint, type checks, and builds. Each produces a result we can inspect: what passed, what failed, and where. Coverage reports help us see what we have not tested. An agent saying the work is finished is not enough.

We automate native experiences in simulators and emulators, too. On Android, wireless debugging lets us install a build and exercise it on a real phone. These checks bring us closer to the experience a collector will actually have, including the camera, touch targets, and navigation.

Browser tests can also be flaky. We keep their attempt history, look for the same test failing and passing without a code change, and use a managed quarantine process with an owner and a repair deadline. A retry is evidence about instability. It is not permission to forget the first failure.

We use a language model as a judge when the question needs interpretation, such as whether a pressing claim is supported by its source. It works from a rubric and known bad examples. That is another kind of evaluation, alongside the deterministic tests that check our code. Neither can tell us, on its own, whether the product is ready for someone else to use.

The useful question is how much of the work we can trust, and what evidence would change our minds.

A working principle, still being tested

Field note / Reliability

The server said success. The collection stayed out of date.

A large collection exceeded the response limit between our backend and the client. The handler finished, but the app still failed to receive the collection. We changed the delivery path and verified the complete request. That incident sharpened a rule we now use throughout the project: check the result where the user experiences it.

Field note / Feedback

“Deployed” was too early to call a bug fixed.

Our feedback tracker used issue closure to tell a reporter their problem was fixed. A merged pull request or a successful deployment could trigger that message before anyone had reproduced the result. We now keep a fix pending until someone confirms the behavior on the released build. Automation can move the work along; it still needs a truthful finish line.

Field note / Evaluation

A convincing answer needs something to answer to.

Pressing Intelligence can attach a real comment to the wrong edition, overstate a reviewer’s opinion, or rank a weak match too highly. We already keep reference cases and evaluate ranking and claim support. Some checks report results for review rather than blocking a release. The broader AI Evaluation Lab is the next experiment: test the models, prompts, and handoffs themselves.

Read the evaluation design

The things underneath

A real product
on a familiar stack.

We use established tools so the experiment can focus on how we work. Each native app has its own interface; the product rules and data meet in the shared API.

In the browser
React · TypeScript · Vite · Tailwind
On the phone
Swift & SwiftUI · Kotlin & Jetpack Compose
Behind the app
Node.js · API Gateway · AWS Lambda · DynamoDB
Delivery & quality
S3 · CloudFront · GitHub Actions · Vitest · Playwright
Data & AI
Discogs · Claude · Codex · Gemini · native recognition tools

The work continues
between our sessions.

A collector reports a bug. Another opens an album we have not researched well enough. Both can become the next piece of work, without waiting for me to sit down at the desk.

A bug report can reach production before I read it.

It still cannot be called fixed until a person has seen it working. We learned that one the embarrassing way.

T+0

Someone tells us

A note from the app becomes a tracked issue.

On the hour

Triage

Priority, effort, and whether it is safe to fix without a person.

Same run

The fix

A scoped change on its own branch, with the tests that prove it.

~20 min

Checked and live

Review gate, CI, merge, deploy, verify production.

Then it waits

Fix pending

Here the machine runs out of authority. The reporter sees “Fix Pending”.

Only a person

Fixed

Somebody sees the symptom gone on the build the reporter is running.

What it will not touch
Sign-in and sessions, anything holding a token, data schemas, infrastructure and the workflows that enforce these rules. Those paths are protected by the same checks that block a human in a hurry. Ambiguous asks and anything that looks like a redesign stop and wait for a decision.
What happens on a quiet hour
Nothing. The job reads the queue, finds no real event, and exits. Most hours are quiet hours, and a loop that invents work to justify itself is worse than one that goes back to sleep.

Half an hour later, a second job looks at the albums collectors actually opened and finds the gaps in our pressing research. Demand tells us where the work would be useful; our evidence rules decide whether it belongs in the product. Sometimes the right outcome is to leave a record with facts only, or change nothing.

The loops are deliberately not the same shape. A code fix, an additive data update, a model extraction, and an analytics refresh each need different checks. They have locks, failure reporting, and a way to stop them. They are scheduled software with an agent inside, and they need maintenance like the rest of the product.

Every hour · :00

Listen to collectors.

Feedback triage and permitted fixes.

Every hour · :30

Follow the records.

Research gaps from recent album opens.

Scheduled batches

Keep the evidence useful.

Video extraction, reporting and cost history.

The hourly jobs are installed local scheduled tasks. Cloud schedules are defined in GitHub Actions. The paper documents the current configuration, limits, and differing release paths.

Follow the scheduled workflows

Did we help someone
understand a pressing?

Our North Star

Total Pressing Intelligence
Panels Viewed

Cumulative views of panels that deliver usable pressing information to a collector.

We count the moment a collector receives usable pressing information. That is close to the reason the product exists: helping someone work out which version to own, buy, or keep. Generic page traffic does not tell us whether we delivered that value.

We work on both sides of that moment. Search and scanning help collectors find the record. Collection and wantlist browsing put their own music in reach. Stories and Collectors Corner give them reasons to explore. The PI research loop improves the evidence for the albums they open.

A growing counter needs questions around it. Are different collectors reaching PI, or only the same few? Is useful information available when they ask? Do their collections load? Does the panel take too long?

40%PI activation
80%PI availability
95%Collection load success
< 2sPI P90 load time

These are our guardrail targets, not a claim about today’s results. P90 means nine out of ten measured loads finish within that time. Definitions, sample sizes, and coverage matter.

The Admin dashboard gives me the detail behind those questions: daily, weekly, and monthly use; return visits; platform and version adoption; collection failures; search and scan outcomes; source clicks; wantlist and collection actions; and feedback. It also shows how people explore corners, open stories, choose themes, sort, filter, and roll the dice for a record.

Reliability and economics are part of the same view. We track errors, latency, rate-limit pressure, modeled usage costs, and actual AWS cost history. Estimates and bills have different labels. Missing history stays visible. I want to understand why the number moved before deciding what to build next.

See the metric definitions and telemetry

Someone still has
to bring it into the world.

Deadwax has given me a reason to learn the work around the software, too.

That has included preparing store releases, making launch material, testing ads on Meta platforms, and figuring out how to introduce the product to collectors without talking like a software pitch.

Agents help me research options, draft material, and compare approaches. I still have to choose what to say, check the claims, decide what to spend, and put my name behind the result.

Ads are their own product problem.

We ran a small Facebook and Instagram advertising test in August. Ad review also changed the creative: a paid ad needed a different treatment from an organic product demo. For the Android launch, the plan separates store destinations and creative by platform so the results will be easier to interpret.

Meet collectors where they are.

Our September Google Play announcement went to the Steve Hoffman Music Forums. It thanked beta testers, showed the collector workflow, and invited corrections. Caden brings the same attention to our Instagram presence. The app has to earn interest through something people can use.

Measure more than a click.

The marketing question is whether someone reaches the store, tries the app, and finds enough value to return. We are still learning how to measure that journey. We do not have a validated acquisition-cost or advertising-return story to tell yet.

This is what makes Deadwax a useful lab for me: product decisions, engineering choices, costs, trust, and distribution all affect one another. I get to see where my assumptions meet the real world.

A roadmap with
some miles on it.

Android is out of beta. Core search, scanning, and the prioritized Pressing Intelligence expansion have shipped. The next work goes deeper on quality, feedback, and the value of the collection itself.

In collectors’ hands

Available
  • Web, iOS & Android v1

    The app is public on the web, the App Store, and Google Play. Collection and wantlist browsing share the same backend.

  • Search, scans & pressing research

    Forgiving search, native barcode and label/deadwax workflows, and source-backed pressing context. Coverage continues to vary by album.

  • Read the label. Speak the runout.

    Native label capture and spoken deadwax make faint markings searchable. iOS cover recognition adds an album-first entry point.

  • Use the collection you have

    Sorting, filters, current record values, and dice or shake-to-roll help collectors spend time with their own music.

  • The PI evaluation foundation

    Reference cases, output-ranking checks, and a separate claim-support judge. Weekly output evaluations currently report for review.

Working on / up next

Current priorities
  • AI Evaluation LabFirst priority · baseline phase

    Establish baselines, then compare models, reasoning effort, prompts, and workflows on fixed tasks. Measure correctness, cost, time, and review burden.

  • Stories & Collectors CornerFinish native distribution

    Album Stories and 18 curated corners connect records to their scenes. Already live on web, the latest native updates are built and tested. The iOS build is uploaded to App Store Connect; Android is signed and device-tested. The next step is completing human release checks and public store distribution.

  • Ask for feedback at useful momentsNext product priority · planned

    Invite feedback after meaningful collector workflows, with careful triggers and limits so the app does not become an interruption.

Further out

Planned / research
  • Collection value & insights

    Build on current record values with better collection insight. Historical and market-source work depends on the data we can responsibly use.

  • Better handoffs between agents

    Evaluate what survives a change of model or working environment: decisions, constraints, unfinished work, and the evidence behind them.

  • Stereo equipment planning

    Explore useful gear context and AI-assisted discussion. Privacy, sources, cost, and scope need decisions before this becomes a product promise.

These are priorities, not release dates. We update the order as we learn. The technical paper separates the evaluation tools that exist today from the lab we are planning.