Ship or Shut Up
AI builds 10 products. Will any of them sell?
Updated nightly from my own log. Last update 23 Aug.
The bet
I have built millions in revenue. For other people.
Eleven years of it. Payment infrastructure that moves billions a year. Mobile apps in production, platforms for 80+ companies, over 100 engineers trained. My name is on none of it, and none of it is mine. The competence is settled. The ownership is not.
And before you assume I never tried on my own: I have shipped products. They are live right now. They work. They made nothing.
I know exactly why, because it is the same three things every time. I cannot call anything finished, so it sits for weeks. When it finally ships, I barely market it. Then a new idea shows up, and building is fun while selling is not, so I go build. The first one costs weeks. The other two cost everything, and they are the same problem twice over.
That is not a skill gap. I could learn the skill in a month. It is that I do not want to do it, and wanting is harder to fix than knowing. So for 30 days I am doing the opposite, in public, with the numbers on this page.
Up to 10 products. AI writes every line, I direct it. $100 total budget for the month. Target is $1,000 in revenue or $100 MRR by 31 August. Success is one product that makes money, not ten that ship. Every number here updates daily, zeros included, and right now they are all zeros.
What the $100 does and does not cover
It covers domains, ads, and any new tool the sprint needs. It does not cover Claude at $100 a month and ChatGPT at $20 a month, which were already running before day 1 and keep running after day 30. Flagging it because the constraint is the entire point: more money makes making money easier, so the test only means something if the incremental spend stays small. Check my arithmetic rather than taking my word for it.
What this sprint is not testing
Whether I can build. That question is closed, and the receipts are not anonymous.
Every one of those numbers was earned for somebody else. So the variable being tested here is not whether the thing gets built. It is whether I can sell it once it exists, which is the part I have failed at every single time. That is why a $0 on this page is the experiment working, not the experiment failing.
The rules
Written down before day 1, so they cannot be renegotiated
Every rule targets one specific way my previous products died. None of them are arbitrary, and none of them were written after the fact. If I skip one, the honest thing is to name which failure just got left unguarded.
2 to 3 days per product. Hard cap.
Guards the polish spiral
I cannot call anything finished, so things sit for weeks while I add what nobody asked for. A hard cap means the deadline decides, not my taste.
One full day of actually selling each product.
Guards the marketing gap
Not the launch post. A whole day of pushing it. This is the rule that carries the entire experiment, because this is where every previous product died.
That selling day happens before the next build starts.
Guards the jump
Building is fun and selling is not, so a new idea always shows up right when the pushing should begin. Ordering the rule this way is the only thing that stops it.
Pre-build gate. Two of three must pass.
Guards against graveyard ideas
Can I point at three places people are already complaining about this? Is someone already charging money for a fix? Can I reach those people? Fewer than two and I skip it.
Stick signal, written down before launch.
Guards against moving the goalposts
One paying customer within 72 hours, or 30 signups within 7 days. Decided in advance it is a test. Decided afterward it is an excuse.
Day 15 checkpoint.
Guards against busywork
If any product has revenue, the build queue stops and the rest of the sprint markets that one. Shipping a sixth product at $0 would feel productive and would be the old pattern wearing a sprint t-shirt.
What is live
The products
The log
Every day, including the bad ones
Build hours and selling hours are counted separately on purpose. If the selling column drifts toward zero while the product count climbs, the sprint is reproducing my old failure at higher speed, and that is worth more to anyone reading than another launch.
Day 12. My ads have now delivered zero impressions three days running and I did not notice until a scheduled check told me. Not a bad result, no result. €1 a day against a tier-1 audience with two placements does not clear Meta's delivery threshold at all, so five of my seven test days are gone and I have learned nothing about the creative.
Then I spent the morning building something else. A brand new faceless horror channel, from an empty folder to a brand doc, an avatar and a working pipeline, on day 12 of a sprint whose entire point is that I stop doing exactly this. Product one is at 0 signups of 30 with two days left on the window and it got zero minutes of my time today.
I have written a rule against jumping to the next build. It only guards the ten products on the list, and today's build was not on the list.
Day 11. I finally looked at the Meta ads I have been running since day 7 and they were buying the cheapest impressions on earth. 71% of delivery went to Syria, 2 impressions reached the United States, and Syria cannot pay Stripe. €3.85 bought 2,236 landing page views that were never able to become a customer.
Then it got worse. Google Analytics recorded 77 real sessions against Meta's 2,236, a 29x gap. And when I filtered to the countries I actually advertise to, the real visitors spent 14 seconds on the page and looked at 1.06 pages. Twice today I found a lead in the data and twice it turned out to be me, browsing my own admin panel and polluting my own numbers.
Rebuilt the campaign for US, UK, Canada and Australia, rewrote the landing page so it says the same thing the ad says, and kept the budget at €1 a day. Day 11 of 30, still 0 of 30 signups, still $0 in.
Day 9 went to my software company. Three to four hours of meetings on client projects and sales, ending at nine at night, which is the whole focus window this sprint gets on a normal day. Nothing was built and nothing was sold.
That's the part of a 30 day sprint nobody puts in the plan. I budgeted one to two focus hours a day and never wrote down what happens when the business that funds it needs the same hours. Day 9 of 30, still 0 of 30 signups, still $0.
Nothing posted, no signup, and the day still went on selling. What it produced was a list.
I scanned 331 domains off one launch directory looking for a specific person: owns a site, is plausibly paying for things already, and has a blog that died months ago. 276 were reachable. 220 had a pricing page. 85 had a blog abandoned for 90 days or more, or no blog at all. Of those, 36 had an X account I could actually find. Roughly one in nine survives the whole funnel.
Two things about that list need saying plainly. Revenue was never verified, because nothing in a public scan reveals it. A pricing page was used as a proxy for "can pay" and it is labelled as a proxy, not as a fact. And about one in four entries was noise: one domain's blog index turned out to be a shared template, and five unrelated domains carried the identical date on it.
The targeting started wrong and it was my error. My first instinct was to message every builder on X with a product, which is close to the worst possible target for a writing tool, because writing their own posts is the entire thing those people are already doing. The corrected version is one line: find people making money who do not have a functioning blog.
I also killed an outreach tool that would have sent a thousand emails on behalf of a domain that was four days old. Their own published numbers are a thousand emails for three to five replies, and the consent line asks permission to email as me. A domain that young cannot afford to have its reputation spent that way.
Twelve messages written, one per prospect, each differing in the observation rather than in a merge field. Five prospects deliberately skipped, including one that turned out to be an Arabic-language platform. Sending to that one anyway is exactly what makes a batch read as a blast. None of the twelve have gone out yet.
Fifteen deploys between midnight and six in the morning, and every one of them fixed a measurement rather than a feature. The question I kept asking in different shapes was whether the thing I was measuring was the thing I cared about. The answer was no six separate times.
The one that turned the night: an article scored 87 out of 100 and passed my quality gate. Then I read it. Four customer stories in it about customers who do not exist, and eight sections describing the same seven tools. Content Depth scored that article 100, because depth was counting words, and both repetition and invented case studies produce words.
Then the sourcing score, which turned out to be inverted. Articles carrying at least one unlinked attribution averaged 90.4. Articles with clean sourcing averaged 70.7. The reason was in my own arithmetic: any external link earned 5 points up to a cap of 25, and a statistic confirmed against the page it came from earned 15. Decorating a paragraph paid better than sourcing it, so that is what the model did. It found the maths before I did.
The rest of the day went to a single Reddit post. Nine rejected drafts and about five and a half hours, which is the worst ratio of effort to outcome in the sprint so far. It got 468 views and 2 upvotes.
The fixable part of that was mine. The post argues from a screenshot, and I put the screenshot in a comment instead of in the post. 468 people read the sentence pointing at the evidence. 13 of them saw the evidence.
One thing got turned down. Faking the screenshot came up as an option and was killed. What replaced it was rendering the real panel from real rows, so the pixels were generated here and every number in them is genuine. Every rule on this page depends on the numbers being real, and that post's whole argument is that my figures survive being checked.
Zero signups, and about an hour of day 4 spent seriously considering dropping product one to go and build something that sells more easily.
What stopped that was checking what I had actually done to sell it. Two posts had gone out across four platforms. Neither carried a link. Roughly 1,090 people saw two days of my content and not one of them was told what I had built. So zero signups was not the market saying no. There was nobody in the room to say it.
I am fairly sure that is how my previous products died. Not from no demand. From no distribution, written down afterwards as no demand. Those two look identical on a dashboard and only one of them is a reason to stop. I was about an hour from reading an empty measurement as a verdict and going off to start the next thing, which is the exact move this whole sprint exists to catch.
So the rest of the day went on the product that already exists instead of the next one, and the content queue changed shape. Three failure stories were written and ready and none of them was a launch. Every batch can carry a link now, and the launch post is the only piece of content that can move the number I registered.
Also recorded, because burying it would defeat the point of this page: I moved the signup deadline from 12 August to 15 August, roughly two hours after the product was verified live from outside. There were zero signups on the board at the time, so there was no result to work backwards from. The old date stays written down next to the new one.
Product one went live on day 3, on a domain I did not own that morning. $11.48 for heybyline.com at 19:59, pointed at the product six minutes later, then a deploy pipeline and nine commits bootstrapping the server it runs on. It was serving by 23:33. Stripe went from half wired to shippable in the afternoon, with the approval path written down end to end. Nobody was told, and none of the numbers on this page move because a URL exists.
The day was 9.4 hours around a fourteen hour hole. I stopped at 01:48 and came back at 16:08, four hours later than I said I would. The work in this note sits either side of that gap.
The fix worth the whole day was in a code sample. The webhook example on Byline's settings page, the snippet a customer copies into their own production site, skipped signature verification whenever the header was absent. Leave the header off and nothing is checked. It also verified the wrong signature, a body only digest with nothing in it that expires, so one captured request replays forever. I found it for a stupid reason: I had fixed the identical fail-open shape in a cron route the day before, so the pattern was still in my head when I read past it. It lives inside a string in a page component, which means no test, no type check and no linter in that project could ever have caught it. It is not code that runs in my repo. It runs on other people's servers.
Three things failed silently that day, or would have. Autopilot is an hourly cron and the failure mode of a cron is silence, so it writes a heartbeat now: the health endpoint reports stale after two and a half hours, and Sentry fires only after two failures in a row. The first version of the deploy skipped itself quietly when its secrets were missing, and I changed it to fail loudly two minutes later, in the same block it was written in. And the script that publishes these notes had been reading the first note in a day file and stopping, exiting clean, doing half the job. Nothing reported any of the three. A job that runs, exits zero and quietly does half the work looks exactly like one that works.
Two more bugs turned up in the output itself, found by reading real articles rather than by any check that was supposed to catch them. One had a heading fused onto the end of a sentence and rendered a literal hash pair in the middle of a paragraph, and it passed the quality gate at 87 out of 100 with the markup visible in its own preview. Twenty two content versions carry words from my own AI tell blacklist, and the gate scored brand voice on every one of them and passed. Third proof in two days that the thing doing the measuring is not measuring what it claims to.
For a full day my own log said I had bought a different domain. The message I sent named one, a commit six minutes later pointed the entire codebase at another, and neither of the two sessions running that evening could see the other one. It surfaced only when the two records were put side by side. The wrong commit message stays as it is, because editing a record to hide that it was wrong costs more than carrying the correction does.
That blind spot is the real finding. Nineteen of the twenty five commits I made on day 3 were missing from this log when I first wrote it up, because two sessions ran in parallel, I moved between them, and each recorded only what it could see. The hours are restated because of it: 0.8 deciding, 5.4 building, 3.2 selling. That is an estimate rather than a measurement, since two tracks running at once overlap in wall clock instead of adding.
Money: $11.48 on the domain, $0.48 of API spend on a real generation run. $17.96 of the $100 gone by day 3, eighteen percent of the budget, none of it on anything a customer can see yet.
Day 2 opened at three minutes past midnight, still inside day 1's session. One thing left broken: every article the engine produced had zero internal links. Turned out the code was working correctly and the input was garbage. It was being offered the site's own privacy policy as a link target, and the AI editor kept refusing to link it, which was the right call. The fix that actually worked came from a question, not from more code.
Then a test suite got written and found three more bugs that every check before it had missed, including one that had been quietly merging every bullet list in every article into the line above it. Two of the day's failures were the measurement being wrong rather than the product. A test that fails in the strict direction is more dangerous than no test, because it burns hours fixing things that already work.
The scheduled job that publishes these numbers to my site had also never run once. It fired on time every night and failed every time, on a file that existed and was executable. macOS blocks scheduled jobs from reading the Desktop, which is where everything lives. Worth saying plainly: a job that is scheduled, loaded and silently failing looks identical to a job that is working, right up until you read the log.
Later the same day I opened the product on a brand new account for the first time, and spent the rest of it finding out it did not work for anyone who had not already been using it. A new account could not add a website at all, because the option selected by default sent an empty value the server rejected, and the error it returned named no field. There was no way to sign out. The welcome screen greeted new accounts with an invisible headline, white text on cream. The free trial advertised on the pricing page did not exist in the product, and every new account was being created as a paying customer that had never paid.
I use this thing hard and none of it had ever surfaced, because my own account was already set up properly. The bugs all live in the first ten minutes, which is the only part a returning user never sees again.
The one worth saying out loud: image generation had been broken and nothing said so. No automated test caught it, not one. It only surfaced when I sat down and tested it by hand.
First real money out today. Six dollars, on verifying a product rather than building or selling one.
Twenty one ideas killed in one day, and four hours of the first day went into choosing rather than building. Most died in minutes. One died after I had already written a full spec for it. The one I picked was never on any of the lists, it surfaced from a straight question about a repo I already owned. It is a problem I have been trying to solve for myself for months, and it still took twenty one dead ideas to notice the one I was already living with. Everybody measures the building. Nobody measures the deciding.
Product one is an SEO content engine. Once the direction was locked in, I spent the rest of the day in a deep analysis session with Claude and ChatGPT. We documented every feature, mapped the user flow, defined the generation pipeline, and turned a rough idea into a detailed product specification. That document then went into Claude Design to generate the initial UI and UX. As soon as the designs were ready, I handed everything over to Claude Code to scaffold the entire project, including the architecture, landing page, and the first version of the content generation pipeline.
Day 1 wasn’t about shipping. It was about making sure Day 2 starts with building instead of thinking.
Baseline recorded the night before the sprint. Nothing built, nothing spent, nothing earned. Writing the rules down now so they cannot be renegotiated on day 12 when a new idea looks exciting.
The raw table
The whole thing, one row per day
This is the file I actually keep, rendered straight out. It gains a row a day and it cannot be tidied up retroactively without that being obvious.
| Day | Date | Build | Sell | Shipped | Revenue | Spend |
|---|---|---|---|---|---|---|
| 00 | 1 Aug | — | — | 0 | $0 | $0 |
| 01 | 2 Aug | 2h | 0h | 0 | $0 | $0 |
| 02 | 3 Aug | 10.4h | 4h | 0 | $0 | $6 |
| 03 | 4 Aug | 5.4h | 3.2h | 1 | $0 | $17.96 |
| 04 | 5 Aug | 8.9h | 1h | 0 | $0 | $17.96 |
| 05 | 6 Aug | 5h | 2.9h | 0 | $0 | $17.96 |
| 06 | 7 Aug | 1.1h | 1.4h | 0 | $0 | $17.96 |
| 07 | 8 Aug | 1.1h | 3.5h | 0 | $0 | $17.96 |
| 08 | 9 Aug | 0h | 0h | 0 | $0 | $17.96 |
| 09 | 10 Aug | 0h | 0h | 0 | $0 | $17.96 |
| 10 | 11 Aug | 0h | 0h | 0 | $0 | $17.96 |
| 11 | 12 Aug | 1.3h | 2.4h | 0 | $0 | $0 |
| 12 | 13 Aug | 0h | 0h | 0 | $0 | $0 |
| 13 | 14 Aug | 1h | 0h | 0 | $0 | $0 |
| 14 | 15 Aug | 0h | 0h | 0 | $0 | $0 |
| 15 | 16 Aug | 1h | 0h | 0 | $0 | $0 |
| 16 | 17 Aug | 0.5h | 0h | 0 | $0 | $0 |
| 17 | 18 Aug | 0h | 0h | 0 | $0 | $0 |
| 18 | 19 Aug | 0.8h | 0.1h | 0 | $0 | $0 |
| 19 | 20 Aug | 0h | 0h | 0 | $0 | $0 |
| 20 | 21 Aug | 0h | 0h | 0 | $0 | $0 |
| 21 | 22 Aug | 0h | 0h | 0 | $0 | $0 |
The blueprint
On day 30 I publish exactly what happened
If it works, you get the playbook. If it fails, you get the specific point where it broke and the numbers attached to it. The failure version is not a consolation prize. It is the data point missing from every build-in-public story you have already read, because people whose public builds made nothing tend to go quiet.
Day 30 blueprint
Get the writeup, win or lose
One email on 31 August with the full result: what sold, what did not, the hours, and the point it broke if it broke. No drip sequence in between. Unsubscribe any time.
Or watch it happen daily
Same numbers, posted as they land.
This is my own experiment, run on nights and gaps. If you came here to hire rather than to watch, the client work is a separate thing and it is not run on a $100 budget.
See the client work