AI Products & Strategy September 1, 2026 · 12 min read

When Shipping Is Free: First-Principles Questions for Delight and Bloat

When AI coding tools dissolve engineering scarcity, the passive filter that kept products coherent dissolves with it. This is the set of twelve first-principles questions Product Managers need to develop as the replacement: six to diagnose bloat, six to reveal delight.

By Vikas Pratap Singh
#ai-products #product-management #product-strategy #ai-era #first-principles

The Features Nobody Defends and Nobody Deletes

The last Data Catalog I reviewed at a client had 247 features in the administration panel. Nobody in the room could name twenty of them. The top five accounted for roughly 80 percent of weekly activity. The bottom hundred had not been touched in the tool’s telemetry window. Nobody was proposing to remove any of them, because every one of the bottom hundred had, at some point, been requested by a stakeholder who would notice if it left.

That product is healthy on every dashboard the organization looks at: engagement, NPS, renewal rate, story points delivered. It is also quietly dying of feature accretion, and the leadership team cannot see it because the instruments they have do not measure feature accretion. They measure motion.

This is the problem that Product Management is about to inherit at an accelerated pace. On February 19, 2026, Boris Cherny, Creator and Head of Claude Code at Anthropic, told Lenny Rachitsky that coding is “solved.” He says he no longer manually writes code, relying on Claude Code instead. Inside Anthropic, pull-request throughput rose about 67 percent even as the engineering team doubled, a gain the company credits to Claude Code. By January 2026, 74 percent of developers worldwide had adopted specialized AI coding tools.

When the cost of building drops to near zero, the passive filter that used to do most of a Product Manager’s rationalization work vanishes. The question for every Product leader this year is not “how do we ship more?” It is: “what principles filter what we ship, now that engineering scarcity no longer does it for us?”

Why Questions, Not Frameworks

In Part 1 of this series, I made the engineering-side version of this argument. AI coding tools are expanding systems without improving them because the constraint that produced quality (human time) has been removed. The studies cited there back the pattern.

The product-side version has attracted a full shelf of commentary already. Portfolio rationalization, subtraction audits, feature-kill frameworks, taste-at-speed workflows. They are not wrong, but they all have the same shape: take a playbook from another domain, port it to software, prescribe a process. What they rarely do is hand the reader a set of questions they can sit with in their own product.

That is what is missing. A framework tells you what to do. A question teaches you what to see. In an era where the shape of a product changes faster than any quarterly playbook can respond to, the durable skill is the ability to look at any feature and apply a small number of honest questions that reveal whether it belongs.

The rest of this article is those questions. Twelve of them. Six diagnose bloat; they are mostly answerable with existing telemetry. Six reveal delight; they require listening, not dashboards. The asymmetry is the point. Bloat is a condition you can measure. Delight is a quality you can only recognize.

Part 1: Six Questions That Diagnose Bloat

Bloat is the accumulation of features that technically exist but no longer earn their maintenance cost. It is the default state of any mature product because adding features is a career-rewarded action and removing them is not. Klotz and colleagues showed in Nature in 2021 that across eight experiments, people overwhelmingly default to additive changes even when subtraction is objectively the better solution. The bias is cognitive, not tactical. It shows up in Lego towers, recipes, essays, and mini-golf courses. The same bias operates in product backlogs.

These six questions are designed to expose features that the organization’s incentives will otherwise protect.

1. The first-day question

Would a brand-new user who signed up today notice this feature exists on their own? If the honest answer is “only if we showed them a tooltip,” the feature has outgrown the product. The tooltip is an apology for complexity. Every tooltip, every empty-state nudge, every feature-discovery banner is evidence that the product is hiding something from itself and then reminding itself where it is.

2. The renewal question

In the last twelve months, how many customers named this feature as a reason they renewed? How many named it as a reason they churned? If both numbers are zero, the feature is creating no observable value and destroying no observable value. That sounds neutral. It is not. A feature that sits at zero still accrues engineering maintenance, documentation surface, support load, onboarding complexity, and interface clutter. Zero in, zero out, with a fixed cost still running. That is a negative position on a time delay.

3. The counterfactual sunset

If this feature did not exist today and a Product Manager pitched it in next Monday’s planning review, would it clear the bar we use for net-new features? If the answer is no, the feature should not be in the product either. The asymmetry most teams feel (it is easier to keep than to build, and easier to keep than to remove) is the subtraction bias at institutional scale. The counterfactual-sunset question is the simplest instrument I know for neutralizing it.

4. The documentation tax

If we deleted this feature tomorrow, how many pages of documentation, onboarding flows, training videos, help-center articles, and support macros would evaporate with it? The size of the evaporating pile is a rough measure of the feature’s local complexity. A feature with disproportionate documentation surface is a feature the rest of the product is working around. The docs are the scar tissue; the feature is what caused the wound.

5. The support fingerprint

Pull the last quarter of support tickets mentioning this feature. Read twenty at random. Do they begin with “how does X work” (normal learning) or with “why does X do Y instead of Z” (confusion by design)? A feature’s support fingerprint reveals whether the underlying design is discoverable or whether users are navigating around a design that does not match how they think. Features with a confusion-by-design fingerprint never resolve; they just generate a tax on the company in perpetuity. The Hick-Hyman Law, established in 1952, explains the mechanism: every added choice increases decision time on every other choice. The support queue is where that tax gets paid.

6. The opportunity cost

What did the team not build while they built this? Would that alternative have been more valuable? This is the most uncomfortable question on the list because the honest answer is almost always yes. But the discomfort is the instrument. Every feature in a product displaced something else. Product Managers who never ask this question never practice the judgment the AI era now requires, because the scarcity of engineering time used to answer it for them and that silence is about to end.

How to build the check. Pick one product surface. Run questions 1 through 6 on the bottom decile of its features first. Expect to be surprised. In every organization I have done this with, the fail rate on question 3 (counterfactual sunset) alone is above 50 percent. The first retirement batch is always the hardest because no one has practiced saying no. By the third batch, the muscle starts to form.

Part 2: Six Questions That Reveal Delight

Delight is not the opposite of bloat. A product can be bloated and still contain pockets of delight; a product can be lean and still produce none. The two signals live on different axes, and they are detected by different instruments.

Bloat is a condition you diagnose. Delight is a quality you recognize. It shows up as unprompted language from users, as sticky retention curves, as substitution costs when a feature is threatened. None of these live in a standard product dashboard. They live in what users say, and what they would do if the feature disappeared, and what they return to without being nudged.

These six questions are instruments for listening, not diagnosing.

7. The unprompted evangelism question

Do users describe this feature to other users without being asked? Not “would they recommend the product as a whole” (NPS); not “did they rate the feature highly when prompted” (CSAT). Do they mention it in a Slack thread, in a sales demo they are running themselves, in a customer-to-customer conversation where your team is not in the room? Spontaneous mention is the clearest signature of delight, because spontaneous mention costs the user something to say and they are saying it anyway.

8. The retention curve

Of the users who try this feature in their first week, what fraction are still using it at week six? Delight has a characteristic curve shape; it flattens high. Novelty has the opposite shape; it spikes and decays. A feature whose six-week retention is trending toward zero is a gift the user politely unwrapped, used once, and put on the shelf. That is not delight. That is decoration.

9. The substitution test

If this feature stopped working tomorrow, what would your users actually do? If the honest answer is “probably nothing, they would be fine,” the feature is not delighting anyone. Delight has a substitution cost. When a feature is genuinely delightful, losing it is felt as a loss; users find workarounds with visible friction, or they complain loudly, or they open tickets within hours. When a feature is not delightful, its absence is felt as a room getting cleaner.

10. The regret question

Ask your twenty most engaged customers a single question: “If the product stopped shipping new features for a year but continued to work exactly as it does today, which existing feature would you most regret if it were removed?” Their answers are your delight A-tier.

Notice the structure of the question. It forces a trade-off. Trade-offs surface real preferences. Satisfaction scores, feature-rating surveys, and NPS all surface polite preferences. The regret question surfaces the preferences a user would act on with their wallet, and that is the only kind of preference worth optimizing for.

For practitioners: Do not delegate the regret-question interviews. Book the calls yourself. The tone of the answer, the pause before the feature is named, the language the customer uses to describe what they would miss, are the actual data. A research-function summary will compress exactly the signal the question is designed to produce. Expect to be surprised by what your best customers name, and expect to be more surprised by what they leave out.

11. The workflow-depth test

Is this feature a step inside a larger sequence the user is trying to complete, or is it a dead-end by itself? Delightful features are almost always embedded in workflows. They compound with the features upstream and downstream of them; the user arrives at the feature from somewhere and leaves for somewhere. Isolated features (one-click this, quick-access that, standalone tool inside a suite) decay in usage because they carry no onward motion. Users need to be going somewhere for a feature to matter, because features are verbs and verbs need subjects and objects.

12. The specificity test

Does this feature solve a problem that a generalist product could also solve, or does it solve a specific problem that only this product can solve? Delight almost always comes from specificity. Generic capabilities, no matter how well executed, feel interchangeable and earn no loyalty. The capabilities that reflect a detailed model of how this particular kind of user does this particular kind of job are the ones that create the emotional binding between a user and a product. When you are deciding which features to protect under AI-accelerated pressure, specificity is the most reliable proxy for durable value.

Using the Questions

The temptation, reading a list like this, is to convert it into a checklist that every feature passes through on the way to a Jira status change. Resist that temptation. A checklist turns a diagnostic instrument into a compliance ritual, and compliance rituals are exactly how bloat accretes in the first place.

Use the questions instead as a habit of attention. Pick three features a week. Walk through the twelve questions as an exercise, not as a gate. Disagree with your own first answers. The goal is to build the judgment, not to produce the artifact.

Over time, the pattern becomes legible. Some features fail most of the bloat questions and pass none of the delight questions; they should be retired, and the only real question is when. Some features pass most of the delight questions and fail none of the bloat questions; they should be protected, invested in, and made more visible. The interesting cases are in the middle, where a feature creates real delight for a small population and also carries real bloat cost for everyone else. Those cases are what Product judgment is actually for. No framework will answer them. A practiced set of questions might.

What Has Changed Since April 2026

This article was drafted in April 2026 and published on September 1. Two developments in between cut against the premise that shipping stays free, and both belong here.

  • Subtraction became a product. On June 4, 2026, Brave shipped Brave Origin, a $59.99 one-time-purchase browser, free on Linux, that removes Leo AI, Rewards, Wallet, VPN, Tor, Talk, Speedreader, and Brave’s own analytics pings, keeping Shields and core Chromium updates. A vendor charging money for fewer features is the clearest market signal so far that accretion has a price customers will pay to escape.
  • The cost constraint is coming back. On August 12, 2026, Fortune reported CIOs and CTOs reining in AI usage as bills rose: Samsara capped usage for non-technical staff, Docusign cut context-token use by roughly half, and Compass set per-engineer token budgets across two model providers.

If token cost re-imposes scarcity, the filter this article asks you to build by principle will partly return by budget. The twelve questions still apply. The budget simply stops being the only reason to ask them.

Do Next

PriorityActionWhy it matters
This weekRun questions 1 through 6 on the bottom decile of features in one product surfaceHighest-yield, lowest-stakes way to feel the diagnosis working
This weekBook three regret-question interviews with your five most engaged customersThe answer reshapes your investment thesis before the next planning cycle
This monthAdd the counterfactual-sunset question to every feature spec before approvalForces subtraction bias to meet resistance at the point of commitment, not six months later
This quarterRead twenty random support tickets per A-tier feature per quarter, yourselfThe fingerprint you read is the judgment you develop; no dashboard substitutes
OngoingRedefine “productive PM” on your team to include features retired, not only features shippedThe metric you reward is the behavior you get; if you only reward shipping, you only get shipping

What Changes When Shipping Is Free

These twelve questions were always true. In the pre-AI era, a Product Manager could get away with not asking most of them because engineering scarcity was answering them implicitly. The backlog was long; the team could only ship so much; the features that cleared the bar tended to be the ones that deserved to. The filter worked in the background, and nobody had to name it.

That filter is dissolving this year. Claude Code, Codex, and Cursor are now composable, capable, and cheap enough that an organization’s feature output can plausibly double within eighteen months with no change in headcount. When that happens, the implicit filter of engineering scarcity is gone, and the only replacement is the explicit judgment of the Product function. The same tools that make adding a feature cheap also make removing or consolidating one cheap; whether a product accretes or stays coherent depends on whether anyone owns subtraction, not on the tooling itself. If that judgment is not developed, the product will accrete features until it collapses under its own documentation weight. The A/B test data is unambiguous: 70 to 90 percent of features tested at Microsoft, Bing, Netflix, and Airbnb fail to move the metric they were designed to improve. Shipping more untested features is not progress. It is higher-velocity noise.

The Product Manager who cannot answer most of these twelve questions about their own product is shipping with their eyes closed. In the pre-AI era, that was survivable. In the post-scarcity-coding era, it is not, because the volume of additions will overwhelm any passive filter in months. Delight and bloat have always been the two outcomes of feature decisions. Engineering scarcity used to make the decision feel rare. It is about to become the most frequent decision a Product leader makes, and the people who practice these questions now will have something to say when it comes.

Sources & References

  1. Boris Cherny on Lenny Rachitsky's podcast: What happens after coding is solved(2026)
  2. Lenny Rachitsky's recap of the Cherny interview(2026)
  3. Klotz et al., 'People systematically overlook subtractive changes' (Nature, 2021)(2021)
  4. HBR: When Subtraction Adds Value(2022)
  5. Kohavi, Online Controlled Experiments at Scale (KDD Keynote, 2015)(2015)
  6. HBR: The Surprising Power of Online Experiments(2017)
  7. Laws of UX: Hick's Law
  8. Pragmatic Engineer: How Claude Code Is Built(2026)
  9. JetBrains Research: Which AI Coding Tools Do Developers Actually Use at Work(2026)

Stay in the loop

Get new articles on data governance, AI, and engineering delivered to your inbox.

No spam. Unsubscribe anytime.