What Is Good Taste in Software? How to Develop It
Good taste in software engineering is a trainable skill: the ability to tacitly judge the quality of software against external standards. That’s the definition I use in my talks, applied to software, where it means telling that an interface or a piece of code is off before you can say why, by judging it against standards that exist outside your own preferences.
At my first job ever, a creative director named Stefanie Luppa handed me a design that required some dangerous stuff: a div had to be centered. So I wrote display: flex; align-items: center, decided it looked good to me and called her over. She looked at it for maybe 300 milliseconds and said “that is not centered”, and I genuinely replied “Stef… do you even code?” followed by “CSS doesn’t lie” lol. CSS lied: the font had some weird line height that added space under the text, and I drew grid lines over the whole thing to make sure, and she’d seen it instantly. But like, how did she do that? And can the rest of us learn to?
I’m Tejas Kumar, an AI Engineer at IBM and a conference speaker, and this post grew out of my talk Frontend after AI: The New UX at Future Frontend 2026 in Espoo, Finland. It’s for developers and designers who keep hearing that taste is the skill that matters now and want to know what it actually is and how to train it. It won’t teach you visual design and it won’t hand you my preferences: it’s about how judgment gets built, with the research linked so you can check me. I checked my own talk against the sources while writing this, and 8 things I said on stage get corrected below, with all of them in one table near the end. If you came for taste skills, the files people give coding agents like Claude Code so they stop building the same generic page, jump to what a taste skill is.
TL;DR
- Taste is a skill you train, and the definitions I’ve read mostly skip the “external standards” part: judging against something outside yourself separates taste from preference and from bias.
- You develop good taste with practice that has fast, clear feedback attached: immerse yourself in examples that come with a measured verdict, compare before you read the standard, take corrective feedback, write your standard down, borrow from other fields and look at everything around the thing you built.
- The best research I know on expert intuition (Kahneman and Klein, 2009) says gut judgment becomes trustworthy under 2 conditions (an environment where cues reliably predict outcomes, and enough practice in it with clear feedback), and that how sure you feel tells you nothing about whether you’re right.
- Frontend is unusually trainable because a lot of its standard is published and measurable in Core Web Vitals and WCAG 2.2, and we’re failing it: WebAIM’s 2026 report found detectable accessibility failures on 95.9% of the top 1,000,000 home pages in February 2026, up from 94.8% a year earlier after 6 straight years of small improvements.
- A model can’t supply the taste by itself, as far as the evidence goes: a taste skill only gives a coding agent the standards in writing, so somebody with trained judgment still has to write them and judge what comes out.
- After AI, part of the frontend moves into agents and MCP Apps (interfaces a server ships into ChatGPT or Claude), where a brand’s taste has to show up inside somebody else’s app.
What is good taste in software engineering?
Good taste in software engineering is a trainable skill: the ability to tacitly judge the quality of software against external standards. Taste gets mixed up with preference, style, bias and compliance, so here’s how they compare:
| Term | What it is | Example |
|---|---|---|
| Taste | Fast, tacit judgment of quality against a standard outside yourself | Seeing in 300 milliseconds that text isn’t optically centered |
| Preference | What you happen to like | Dark mode, tabs, a serif you’re fond of |
| Style | A consistent set of choices a person or a brand makes | A brand’s typography, spacing and tone |
| Bias | Fast, tacit judgment against a standard that only exists in your head | Deciding someone works in tech support from how they look |
| Compliance | Passing the written checks | A page that passes every accessibility check and still feels terrible |
Notice how bias is the same fast judgment as taste, just against a standard that only exists in your head. Compliance uses the same external standards as taste, and it stops at whether the checks pass. Even that takes a person for a lot of accessibility work, since axe-core (an automated accessibility checker) only claims in its own README to find “on average 57% of WCAG issues automatically”.
Plenty of people defined taste before me, starting with Paul Graham, who argued in 2002 that taste can’t be mere preference, because if you get better at design then your old taste was worse: “Poof goes the axiom that taste can’t be wrong.” Linus Torvalds showed TED a linked list in 2016: the version he liked removes an entry with no if statement for the first entry, because “sometimes you can see a problem in a different way and rewrite it so that a special case goes away and becomes the normal case. And that’s good code.” Then he called his own example too small, since good taste “is about really seeing the big patterns and kind of instinctively knowing what’s the right way to do things”. Sean Goedecke’s version is “the ability to adopt the set of engineering values that fit your current project”, and Antonio Agudo built a review routine on top of it and defines taste as “judgment about which tradeoffs fit a problem”. The definition I use adds 3 words to theirs: trainable, tacit and external.
Taste is trainable
People usually say 2 things about taste: “it’s subjective” and “you either have it or you don’t”. The second one is the old question of whether good taste is innate or learned, and in February 2026 Sam Gerstenzang took the innate side outright: taste has “always been an important trait (not a skill)”.
On stage I credit the whole definition to a 1993 paper, Carol Mockros’s “The development of aesthetic experience and judgement” in the journal Poetics, and I say “it’s not our definition, it’s from the research”. That’s too strong: I could only check the paper’s abstract for this post, and what it backs is the trainable part. Mockros grouped adults into 5 levels of visual art expertise, had them judge 3 paintings, and found that “specific experience and training in art has a relatively strong overall impact on aesthetic judgements”. It’s a snapshot of people at different levels and it didn’t follow anyone through training, so the careful reading is that training is strongly associated with better judgment.
It takes a while though: in a 2015 study in PLoS ONE a 30 minute training video made only “a small effect” on how novices rated art. The intuition research below also says that even in good conditions “Talent surely matters”. On stage I went further than that and said nobody’s really born with amazing taste and we can all train it, and the research doesn’t back me that far: it says training works, and it doesn’t promise that everyone gets to great taste.
Taste is tacit
Tacit knowledge is what you know and can’t fully put into words. The term is Michael Polanyi’s, from his line that “we can know more than we can tell”. Tacitly means you have a gut feeling, like “I just know that’s a fake Louis Vuitton bag”.
On stage I illustrated this with a statue that I said was fake, with a museum label that says “This is probably fake”, and I got the details wrong. If you’ve read Malcolm Gladwell’s Blink, it’s the statue the book opens with. It’s the Getty kouros (a kouros is an ancient Greek statue of a standing young man), it belongs to the J. Paul Getty Museum in Los Angeles (I said “the United States Museum of Modern Art”), and I can’t find a label that said “probably fake”: the Getty’s record dates it “about 530 B.C. or modern forgery”. The Getty bought it in 1985 after a geologist concluded the crust on its surface could only have formed over centuries, and then several experts recoiled at first sight. In the Getty’s own published volume from a 1992 colloquium on the statue, Angelos Delivorrias writes about “the intuitive repulsion it arouses in me” and traces it to “the distillate of lived experience”, and John Boardman writes that “Instinct without experience is useless”.
In that same volume Boardman tested his own instinct: he went through pictures of kouroi (the plural) that are accepted as authentic and asked how he’d react to each one if it turned up on the market undocumented. He wrote that “Several failed the test”, so he wouldn’t call the statue a fake on a feeling alone. As far as I can find nobody has proved it fake, and it’s been in storage since at least April 2018.
Taste is judged against external standards
As human beings we make tacit judgments all the time, and sometimes those judgments are based on our own bias and not on anything external: people take one look at me and tacitly decide I work in tech support. A tacit judgment based on internal bias is a problem, and the same fast judgment based on external standards is what I’m calling taste.
It’s the same with music: you’re at a concert, somebody hits a note that’s off-key, and the whole room winces at the same moment. How do we all know? We’re tacitly judging the quality of the singing against an external standard: relative pitch. In the key of C you can play G, F and A minor, but play a G sharp major and you’re going to have a bad time.
Is taste subjective?
Preference is subjective, and taste the way I’m defining it is judgment against standards that exist outside the person judging.
David Hume got here in 1757. In “Of the Standard of Taste” he retells a story from Don Quixote where 2 of Sancho Panza’s kinsmen taste a hogshead of wine, and one says it tastes of leather and the other of iron. Everybody laughs at them until the barrel is emptied, and at the bottom there’s “an old key with a leathern thong tied to it”. Hume says producing general rules of art “is like finding the key with the leathern thong”, and his description of a good critic lines up with the 6 protocols below: “improved by practice, perfected by comparison, and cleared of all prejudice”.
We should be careful here though, because for art in general people still argue about whether an external standard even exists. A 2020 paper in the British Journal of Psychology argues that the traditional idea of judging merit against standards of aesthetic value rests on flawed assumptions. Frontend is unusual because so much of our standard is published, versioned and measurable.
Why is everyone saying taste is the new core skill?
On February 14, 2026, Paul Graham posted a prediction: “In the AI age, taste will become even more important. When anyone can make anything, the big differentiator is what you choose to make.” 2 days later, on February 16, 2026, OpenAI’s president Greg Brockman posted 6 words: “taste is a new core skill”. That post had about 2.9 million views as of September 2026, and people now search for it as “taste is the new core skill” even though he wrote “a”. The Linux kernel maintainer Jens Axboe replied “Nonsense, taste has always been a core skill. It’s just easier to see now.” Business Insider’s write-up of the whole argument asked the obvious question: “What even is taste? The term is slippery.”
Taste in the age of AI is easier to see because generating got cheap. Sundar Pichai said in October 2024 that “more than a quarter of all new code at Google is generated by AI, then reviewed and accepted by engineers”, and by April 2026 he put it at 75%, “AI-generated and approved by engineers”. Look at the second half of both sentences though: the reviewing and approving is still people. Anthropic wrote in 2026 that as Claude wrote more of its code, “human code review has become a new bottleneck”. Meanwhile in Stack Overflow’s 2025 developer survey, 46% of developers said they distrust the accuracy of AI tools against 33% who trust it, and 66% said they’re frustrated by “AI solutions that are almost right, but not quite”.
It’s like when everyone got a camera and people thought there’d be so many photographers, and yet open Instagram: just because everybody has a camera doesn’t mean everyone’s an amazing photographer. We take roughly 2 trillion photos a year now (an industry estimate), and code is going the same way, where pretty much anyone can generate it and few of us can create beautiful, tasteful products.
A lot of what we generate also ends up looking the same, and the internet’s name for that is AI slop. Anthropic explained why in November 2025 with what it calls distributional convergence: “Without direction, Claude samples from this high-probability center”, which back then meant the Inter font, purple gradients on white and minimal animation. It shows up in writing as well: a 2024 study in Science Advances had 293 writers write short stories, 198 of them with the option of AI story ideas, and the AI-assisted stories were individually better and collectively more alike, which the authors call a social dilemma.
So everyone says taste matters, and plenty of people have written about how to get it. Goedecke’s answer to “How do you develop good taste?” is “It’s hard to say, but I’d recommend working on a variety of things, paying close attention to which projects (or which parts of the project) are easy and which parts are hard”, and most of the other advice I’ve read is some version of look at great work, make a lot of stuff and trust your gut, even though there are decades of research on how this kind of judgment gets built.
Can taste be taught?
Taste can be taught, within limits: gut judgment becomes reliable where cues predict outcomes and feedback is fast and clear, and the cues experts use can be written down and taught to novices.
The best research I know on when you can trust a gut feeling is Daniel Kahneman and Gary Klein’s 2009 paper in American Psychologist, “Conditions for Intuitive Expertise: A Failure to Disagree”, and it isn’t about taste at all. The 2 of them had spent their careers on opposite sides of the question, since Kahneman cataloged how intuition fails and Klein studied firefighters and nurses whose intuition saves lives. They agreed on 2 conditions for skilled intuition: “an environment of sufficiently high validity and adequate opportunity to practice the skill”. High validity means cues reliably predict outcomes, and the practice only counts when the feedback is fast and clear.
On stage I summarized this as immersion plus corrective feedback, and the paper also carries a warning I didn’t mention, that feeling sure tells you nothing about being right: “There is no subjective marker that distinguishes correct intuitions from intuitions that are produced by highly imperfect heuristics.” Their replacement test is to ask how valid the environment is and what the judge’s history of learning it looks like, which for us means asking whether somebody’s taste got trained somewhere with real feedback or just got confident.
That’s also why taste is per task and not per person. Kahneman and Klein call it fractionated expertise and say it’s “the rule, not an exception”: you can have a trained eye for spacing and typography and none at all for information architecture, and feel equally sure about both.
A lot of what experts know tacitly can still be put into words and taught. In the same paper, nurses in a neonatal intensive care unit could tell a baby was developing sepsis before the blood tests could, and “When asked, the nurses were at first unable to describe how they made their judgments.” Researchers interviewed them incident by incident and pulled out the cues, “some of which had not yet appeared in the nursing or medical literature”, and those cues became training material. Irving Biederman and Margaret Shiffrar did the same thing with chick sexing (telling male day-old chicks from female ones) in 1987: they wrote down what 1 veteran expert looked for, and novices’ agreement with the experts went from a correlation of .21 to .82 after reading it.
Applying this to software is my interpretation and not theirs: performance and accessibility are high validity with fast feedback, since you can measure them today. “Does this feel premium” has slow and ambiguous feedback, and long-range bets about what users will want in 3 years behave like the stock pickers the paper uses as its example of intuition you shouldn’t trust.
Where 2 standards pull against each other, like a faster code path against a more readable one, no published threshold settles it, and that’s the part of taste Goedecke and Agudo are writing about. By Kahneman and Klein’s conditions it’s probably also the hardest part to train, because the feedback on an architecture call arrives months later, mixed in with everything else that happened.
Frontend already has external standards
At Future Frontend my talk was in the accessibility block, and it’s so cool that we have actual standards: numbers we can optimize against that already improve the tastefulness of what we build.
Core Web Vitals are Google’s 3 published metrics with thresholds for good and poor: Largest Contentful Paint for loading, Interaction to Next Paint for responsiveness and Cumulative Layout Shift for visual stability. Here they are:
| Metric | Good | Poor |
|---|---|---|
| Largest Contentful Paint (LCP), loading | 2.5 seconds or less | Over 4.0 seconds |
| Interaction to Next Paint (INP), responsiveness | 200 milliseconds or less | Over 500 milliseconds |
| Cumulative Layout Shift (CLS), visual stability | 0.1 or less | Over 0.25 |
A site passes when at least 75% of page views hit “good” on each of the 3, measured separately for mobile and desktop. Those page views come from Chrome users, because the Chrome UX Report (the real-user data Google grades sites on) doesn’t see Chrome on iOS or other browsers. INP replaced First Input Delay on March 12, 2024, so these standards get revised in public. In Chrome’s August 2026 data, 55.6% of 18.3 million sites pass all 3, so roughly 44% still fail.
WCAG 2.2 is the W3C’s accessibility standard: a W3C Recommendation published on October 5, 2023 and updated on December 12, 2024, and also an ISO standard (ISO/IEC 40500:2025). It has 3 levels (A, AA and AAA), and at AA it has testable criteria like text contrast of at least 4.5:1, or 3:1 for large text, and pointer targets of at least 24 by 24 CSS pixels, with exceptions such as links inside a sentence. Core Web Vitals are Google’s published thresholds and don’t come from a standards body, but they’re both external to you and me.
There’s money in it too because nobody’s really going to use your thing if it’s distasteful. On stage I said a 100 millisecond improvement in site speed increased conversion rates significantly, and that came from a Google-commissioned Deloitte study that looked at 37 brands’ mobile sites and more than 30 million sessions in late 2019. It found that a 0.1 second speed improvement was associated with 8.4% more conversions on retail sites, 10.1% more on travel sites, and retail customers spending 9.2% more. It’s an observational study, it predates Core Web Vitals, so the speed metrics it tracked aren’t the 3 in the table above, and every page in the journey had to get faster for the effect to show.
Is vibe coding making the web less accessible?
WebAIM’s 2026 report says “vibe coding” is likely part of why detectable accessibility failures went up after 6 straight years of small improvements: 95.9% of the top 1,000,000 home pages had them in February 2026, up from 94.8% a year earlier. The report found an average of 56.1 errors per page, and its guess is that the trends “likely reflect” heavier reliance on third-party frameworks and AI-assisted coding.
How to develop good taste: 6 protocols
You develop good taste with practice that has fast, clear feedback attached: immerse yourself in examples that come with a measured verdict, compare before you read the standard, take corrective feedback, write your standard down, borrow from other fields and look at everything around the thing you built.
Here’s the recipe from my talk with 1 step added (2) and 1 aside from the talk promoted to a step (6), the research underneath each step and a protocol you can run. Where a study backs a step I say what it measured, and where I’m extrapolating I say that too, because almost none of this evidence is about training anybody’s taste in software. It comes from medical students, art novices and college students, plus 2 studies of how code review works, so applying it to interfaces is a bet: these are things I’d practice and I haven’t run them as an experiment.
1. Immerse yourself in the real thing
Stef could see my centering problem in 300 milliseconds because she’d looked at hundreds of thousands of things before mine. On stage I tell the story that people who catch counterfeit money spend almost all their time studying real money, so that a fake feels wrong immediately. I checked it for this post, and the popular story that people who catch counterfeit money only ever study the real thing is about half true.
The true half: the Federal Reserve’s training for bank tellers is titled “Teller Toolkit: A Guide to Identifying Genuine Currency”, and a Secret Service agent teaching a class in 2020 put it this way: “By understanding how to identify genuine U.S. Currency, we can more easily detect counterfeit bills.” A cash supervisor at the Boston Fed describes what that builds: when you touch currency all day “you get a feel for the currency”, and a fake registers “like a gut feeling”.
The folklore half is the version where trainees never touch a fake. The theologian Roger Olson asked a Secret Service agent who trained bank tellers and the agent “laughed at the story”, and that same 2020 class handed out real counterfeits to compare against. I also said on stage that they don’t even trust AI with this, which is wrong. The Federal Reserve verifies deposited cash “note-by-note, on sophisticated processing equipment”, and people only make the final call on the notes the machines reject. Frontend has the same split available: Lighthouse (Google’s page audit tool) and axe can take the first pass in CI (the checks that run on every change), and a trained person judges what they can’t see.
Immersion on its own can backfire though: James Cutting found in 2003 that adults’ preferences among Impressionist paintings tracked how often those images had been reproduced in books, and concluded that “mere exposure helps to maintain an artistic canon”. Looking at a lot of stuff can simply teach you to like what’s common. So immerse yourself in something beautiful like Linear, and in ugly things too, but attach something checkable to every example: a Lighthouse run, an accessibility audit or the standard it meets or breaks.
Protocol 1: blind calls (15 minutes, 3 times a week). Build a deck of 30 to 50 things whose verdict is already measured: pages labeled pass or fail on a Core Web Vital, components labeled pass or fail on a WCAG criterion, pull requests labeled approved or changes requested. For color contrast, screenshot text on its background and label it with the computed ratio. For Core Web Vitals, don’t judge from a load on your own laptop because the verdict comes from other people’s devices: record the page loading with DevTools throttled to Slow 4G and a 4× CPU slowdown, and label the recording with the field verdict from PageSpeed Insights. For pull requests, gh pr list --state all --json title,url,reviewDecision gives you the labels. Look at one, commit to a call within about 10 seconds, reveal the answer and keep a tally. I lifted this from a histopathology module for medical students, where students classified images with immediate feedback and a category was retired after 3 fast correct answers in a row. The whole module took a median of 11.5 to 17.5 minutes, and scores for fast, correct answers were still well above where they started 6 to 7 weeks later. The jump from tissue slides to interfaces is my extrapolation.
2. Compare first, read the standard second
Comparing things side by side beats looking at them one at a time, and the order matters too. A meta-analysis of 57 experiments found that comparing cases beat studying them one by one (an effect size of d = 0.50, which counts as medium), and that it worked better when the principle came after the comparison. The cleanest demonstration is Daniel Schwartz and John Bransford’s 1998 paper “A Time for Telling”. Students who analyzed contrasting cases and then got the lecture made 43.8% of the possible correct predictions a week later. Students who analyzed the cases twice with no lecture made 16.7%, and students who read a summary and then heard the same lecture made 14.6%. The same lecture was worth 3× as much to the people who’d worked through the cases first.
Protocol 2: contrast, then read (30 minutes a week). Take 2 or 3 implementations of the same thing (3 checkout forms, 3 date pickers, 3 empty states), ideally one that passes a standard and one that fails it. Before reading anything, write down every difference you can see. Only then open the relevant WCAG success criterion or web.dev guidance, or ask a senior person for their read, and mark which differences mattered and which you missed. A fun way to start is to play User Inyerface, a game that’s a deliberately hostile sign-up form, once with a timer running and list every convention it breaks, because that list is a first draft of your own standards. (Careful with the spelling: it’s Inyerface, from the Belgian agency Bagaar in 2018, and the “in your face” spelling of the domain is a parked ad page.)
3. Get corrective feedback, and take it
On stage the core recipe had only 2 ingredients, immersion and this one: you need corrective feedback, and you need to be receptive to it.
Feedback isn’t automatically good for you though: Avraham Kluger and Angelo DeNisi’s 1996 meta-analysis covered 607 effect sizes and found that feedback improved performance on average, but “over 1/3” of feedback interventions made it worse. The pattern was that feedback stops working as it pulls attention off the task and onto the self, so “this isn’t centered” is the kind that helps and “you’re not a details person” is the kind that backfires.
For developers, code review is a feedback loop we already have, and it’s fast and it’s about the work. When Alberto Bacchelli and Christian Bird studied review at Microsoft in 2013, the most common kind of comment was code improvements at 29%, with actual defects at 14%, and at Google in 2018 the median review came back in under 4 hours.
Protocol 3: predict the review (10 minutes per pull request). Before you request review, write down the 3 comments you expect to get. Afterward, score yourself: every comment you didn’t predict is probably a standard you don’t hold yet. The reverse works too, where you review a teammate’s change privately and then compare with what the senior reviewer wrote. When you ask for feedback outside code review, swap “is this good?” for “what would you change first, and which standard is it failing?” so the answer comes back about the task. The prediction step is my addition, since none of those papers tested it.
4. Write the standard down
On stage I said Dieter Rams, Braun’s head of design from 1961 to 1995, advises people to write down what makes good things good. I went back to the sources and that’s not quite right, since all I can find is Rams writing his own standard down. Vitsoe tells it like this: in the late 1970s Rams “asked himself an important question: is my design good design?”, and his answer became the 10 principles of good design. The design historian Klaus Klemp traced the drafts through Rams’s archive and puts the start earlier, with the whole thing taking a decade: 3 rules in a 1975 lecture, 6 principles in 1983, 10 by 1985. Rams treated it as a living document as well, telling the filmmaker Gary Hustwit: “I didn’t intend these 10 points to be set in stone forever. They were actually meant to mutate with time and to change.”
Jony Ive, who led design at Apple until 2019, took it to heart too: in a 2024 interview he said “That is what I and so many designers owe Dieter. He articulated the way it should be”, so if you work on a Mac there’s some Dieter Rams in it by way of Jony Ive.
You can also write down the wrong standard: a review of 75 studies of scoring rubrics found they make judgments more consistent, especially with examples attached, and also that “rubrics do not facilitate valid judgment of performance assessments per se”. A rubric can make you consistently wrong, so anchor yours to standards that exist outside your team.
Protocol 4: the standards file (20 minutes every Friday). Keep 1 file in your repo where each entry has a criterion, a passing example, a failing example and the external standard it traces to. Every miss from protocols 1 to 3 becomes a new or revised entry. Once a month, score 3 old pieces of your own work against it and see whether your scores hold. Here’s what 1 entry looks like, with contrast ratios computed from the WCAG formula:
## Body text contrast
- Criterion: body text is readable against its background
- Passes: #1a1a1a on #ffffff (17.40:1)
- Fails: #999999 on #ffffff (2.85:1)
- Traces to: WCAG 2.2 success criterion 1.4.3, at least 4.5:1
The failing gray comes in at 2.85:1 against a 4.5:1 minimum, and #767676 is the lightest pure gray that passes on white, at 4.54:1. The same file is most of a taste skill for a coding agent, and the section on taste skills covers how to hand it over.
5. Cross-pollinate
You probably know the Steve Jobs calligraphy story from his 2005 Stanford commencement speech: he’d dropped out of Reed College, dropped in on a calligraphy class, and “ten years later, when we were designing the first Macintosh computer, it all came back to me”. There’s a more direct line in a 1995 interview that aired on PBS in 1996, and the transcript has him saying: “Ultimately it comes down to taste. It comes down to trying to expose yourself to the best things that humans have done and then try to bring those things in to what you’re doing.”
This is the least supported part of the recipe because the research on borrowing from distant fields measures how novel people’s ideas are and says nothing about whether anyone’s judgment got better. Nikolaus Franke, Marion Poetz and Martin Schreier found in 2014 that problem solvers from analogous markets produced solutions with “substantially higher levels of novelty” and “lower potential for immediate use”, so expect ideas that need adapting before you can use them.
Protocol 5: borrow 1 idea from a distant field (45 minutes, every other week). Pick a field with a problem shaped like yours (print typography for readability, wayfinding for navigation, film editing for pacing), pull out 1 principle, and write the mapping out explicitly: what in their world corresponds to what in yours. Try it on something small and run it past the same standards and reviewers as everything else.
6. Look around the button
Taste doesn’t just judge better, it perceives differently: if I bring someone early in their career into a room and ask about the call to action button, they’ll tell me it’s blue and it represents the brand well. If I ask somebody more experienced, they’ll comment on the button and then on the things around it: on some devices it renders under the fold, and on slow networks it isn’t visible at all.
Protocol 6: the around-the-button pass (10 minutes per feature). Before you call something done, look at everything that isn’t the component: a throttled connection, a 320 pixel wide viewport (the width WCAG’s reflow criterion uses), keyboard only, a screen reader, 200% zoom, the error state and the empty state. For accessibility, WebAIM’s 2026 data says 6 kinds of error account for 96% of everything it detects: low contrast text, missing alt text, missing form labels, empty links, empty buttons and a missing document language. Pro tip: fix those 6 first, since it’s probably the cheapest accessibility win you’ll get. For speed, use the lab for the fast feedback (a throttled Lighthouse run takes seconds) and the field for the verdict: after it ships, judge on real users at the 75th percentile because your own laptop and connection are probably faster than theirs. The list is mine and not from a study, so add to it every time production teaches you something.
The 6 protocols at a glance
Here’s all of it in one table, with the evidence and my guesswork in separate columns:
| Protocol and time | What the research shows | What’s my extrapolation |
|---|---|---|
| 1. Blind calls (15 min, 3Ă— a week) | Short classification trials with instant feedback built fast, lasting recognition in medical students | That it transfers to interfaces and code |
| 2. Contrast, then read (30 min a week) | Comparing cases beats single cases (d = 0.50), and explanation lands best after noticing | Using good and bad pairs of interfaces |
| 3. Predict the review (10 min per pull request) | Code improvements are the most common review comment (29%), reviews come back within hours, and feedback aimed at the task helps | The prediction step |
| 4. The standards file (20 min every Friday) | Written criteria with examples make judgment more consistent and don’t make it valid | A team file in the repo, reused by an agent |
| 5. Borrow from a distant field (45 min, every other week) | Distant fields yield more novel ideas that are harder to use right away | That it trains judgment at all |
| 6. Around the button (10 min per feature) | 6 error types are 96% of detected accessibility failures | The checklist itself |
The scheduled ones add up to about 2 hours a week, plus 10 minutes per pull request and per feature. The first 3 have feedback built in, so start there.
A live taste exercise: open any website
Besides Core Web Vitals and WCAG, there’s another measure that I personally optimize for: time to done. The truth is I don’t really like using computers because the more time I spend on one, the less time I have with my friends and family. Computers to me are fun to build on and a distraction to do work on, and I don’t like spending time booking flights, I like going on flights.
So in the talk I did an exercise in taste live: open literally any website and feel the pain together. Expedia asked me to sign in before I even knew what it was offering me. On Adidas nothing had happened yet and I was already being asked to do stuff. I didn’t know Coca-Cola’s domain so I Googled it, and got Google in German. When you open a browser and go to some domain, you’re leaving your house and going to someone else’s house, where they’ve made decisions and decorations on your behalf, but they weren’t thinking about you: this is for some legal person somewhere.
Then I played User Inyerface on stage: a timer, fields where the gray hint text is a real value and not a placeholder, so you have to backspace over it, a password that needs at least 10 characters, and a terms checkbox that means the opposite of what you’d expect. It’s the amplified version, but it’s honestly not that far off from regular web use.
Then I showed some good stuff on the web, which was the conference’s own site. Honestly, I love this website: there’s not a single cookie banner on futurefrontend.com, and you go to the schedule and you just see things. I do this talk at other conferences and at this point in it I’m usually totally destroying the conference’s website because their websites usually suck also, but Future Frontend is genuinely good.
The cookie banner thing is measurable, btw. A study from the CHI 2020 conference scraped 680 UK sites that used the 5 most popular consent platforms and found that only 11.8% met even the minimal requirements it derived from European law.
Try it yourself: open 3 sites you’ve never used, start a timer, and write down everything each one asks of you before it has offered you anything.
After AI, part of the frontend moves into agents
So what if I didn’t have to go to somebody else’s house at all, and the data came to me, in an environment I know and trust? That’s kind of what agents allow for. On stage I asked a little local agent I wrote in Node.js (a model inside an agent harness) when the speaker dinner was and how long it would take from the hotel. While I kept talking it searched my email, found the dinner (which, it turned out, had been the day before), opened Google Maps and came back with the hotel in Espoo, the dinner in the center of Helsinki and when to leave. This is what I wanted: I didn’t want to navigate Google Maps with some cookie banner, I just wanted to ask and get an answer.
This runs on the Model Context Protocol (MCP), which works like the web does: a host (an AI app like ChatGPT or Claude) opens a client connection to a server and asks for context (tools, resources or prompts) where a browser would ask for a page. MCP was created at Anthropic and open-sourced in November 2024, and donated to the Linux Foundation’s Agentic AI Foundation on December 9, 2025. I built a tiny MCP server for the conference and asked Claude who was speaking, and on the second try it worked, and it looked kind of nice, but it was kind of soulless. Future Frontend did real work on a logo and a color scheme and we just lose all of it. Imagine Coca-Cola without the red: it’s not Coca-Cola anymore, it’s ChatGPT Cola.
What are MCP Apps?
MCP Apps is the first official MCP extension: an MCP server ships its own interactive interface, and the host (an AI app like ChatGPT, Claude or VS Code) renders it in a sandbox inside the conversation. It went live on January 26, 2026, and the announcement has the mechanics: the host fetches the interface as a resource, renders it in a sandboxed iframe and talks to it with JSON-RPC over postMessage. It’s the fix for the soulless answer above, because the design team’s taste shows up inside the AI app. OpenAI’s docs describe it as “build your UI once and run it across MCP Apps-compatible hosts”. When I added my server to ChatGPT and asked who was speaking, I got speaker cards with avatars and the Future Frontend logo, branded and personal, and there was also text, so it’s accessible.
I was too quick with “so it’s accessible”, because the spec does require the text (“Tools MUST return meaningful content array even when UI is available”), but it’s there for the model and for text-only hosts, and it says nothing about the interface you render. Liad Yosef, one of the co-authors of MCP Apps, put it plainly in an essay with Phillip Lamb: “The rendered view must still be keyboard- and screen-reader accessible, and that conformance is real work that doesn’t go away.”
The upside of moving the frontend into a host is that we say yes or no to cookies once, on ChatGPT, and that’s it. Its accessibility decisions affect everything rendered inside it, which is also the catch: should we consolidate power like that into 1 company? Probably not. The good news is that the protocol is open, so you can make your own client and render beautiful things.
3 things changed between the talk in June 2026 and this post in September:
- July 9, 2026: OpenAI replaced its App Directory with a Plugin Directory, and its developer docs now live at developers.openai.com/plugins, where the old Apps SDK URLs redirect.
- July 28, 2026: a new version of the MCP specification shipped with a stateless core.
- September 8, 2026: version 2.0.0 of the MCP Apps SDK shipped, and the wire protocol didn’t change.
Can AI have taste?
AI can’t have taste by itself, as far as the evidence goes: without direction a model samples the most common choices in its training data, and a written standard only gives it something to be judged against, so a person still has to do the judging. Not everybody agrees: Nan Yu, then head of product at Linear, posted during the February argument that “I hate to break this to everyone, but you probably don’t have better taste than the AI.”
Most of the public evidence is about what a model generates, and taste the way I’ve defined it is about judging. The one experiment I found where a model does the judging is Justin Wetch’s from January 2026, where Claude Opus 4.5 picked between screenshots of generated pages, and nobody checked its verdicts against trained people. When the authors of SkillsBench, a benchmark of 87 tasks that agents attempt with and without skills, had 3 agent setups write their own skills before solving tasks, the self-written skills scored 8.1 to 11.5 points below using no skills at all, while curated skills added 18.2 to 24.8 points on the same setups.
What is a taste skill?
A taste skill is a markdown file of written design standards that an AI coding agent loads before it builds UI, so that its output can be judged against them. It’s what Rams did, written for a model, and it only gives the agent the standards in writing, so somebody still has to judge what comes out. The best-known ones are Anthropic’s frontend-design, Leon Lin’s taste-skill and Paul Bakaus’s Impeccable, and frontend’s published standards are already inside them: as of September 2026 taste-skill’s checklist asks whether Core Web Vitals are “plausibly hit” and requires WCAG AA contrast.
The standards file from protocol 4 is most of a taste skill already. To hand it to Claude Code, give it a name and a description in frontmatter, save it as .claude/skills/<name>/SKILL.md and point to it from CLAUDE.md, because a file only helps if the agent loads it: in Vercel’s agent evals from January 2026, “In 56% of eval cases, the skill was never invoked”. The evidence that these files improve design is thin so far, and I couldn’t find a controlled study with human raters that compares a design skill against no skill.
What standards can’t do
The strongest objection I’ve found to everything above is that taste can’t be reduced to checks. Karl Tryggvason’s essay “You can’t unit test for taste” puts it as “there are no red/green unit tests for taste”. Scott Young, in one of the few pieces on learning taste that cite research, argues that “The rules that guide it can’t easily be written down” and that developing it is “more of a process of enculturation than one of training”. They’re right that the checks don’t get you all the way, because a page can pass every Core Web Vital and meet WCAG AA and still be tasteless. Rams said the same about his own list: “Very good design, however, does not evolve only by ticking those ten boxes.”
On stage I called these standards “objective measures of taste among other things”, and the people writing taste skills call them the floor. Anthropic’s current frontend-design file tells the model to “Build to a quality floor without announcing it: responsive down to mobile, visible keyboard focus, reduced motion respected, visually accessible, harmonious color palettes.” Impeccable’s reference file, which started from Anthropic’s skill, says “The floor holds the mechanics; it never picks the direction.” The floor is also the part you can actually get fast, unambiguous feedback on: contrast, layout shift and input delay.
What I corrected from the talk
Here’s what I said on stage that the sources didn’t back, in one place, with where each one gets fixed above:
| On stage I said | What the sources say | Where |
|---|---|---|
| The statue is in “the United States Museum of Modern Art” and its label says “This is probably fake” | It’s the Getty kouros at the J. Paul Getty Museum in Los Angeles, and the Getty’s record dates it “about 530 B.C. or modern forgery” | Taste is tacit |
| The definition of taste comes from a 1993 paper | The paper’s abstract backs the trainable part | Taste is trainable |
| Nobody’s born with amazing taste and we can all train it | Training works, talent still matters, and the research doesn’t promise that everyone gets to great taste | Taste is trainable |
| People who catch counterfeit money spend almost all their time studying real money | About half true: the training is built around genuine currency, and the part where trainees never touch a fake is folklore | Protocol 1 |
| They don’t even trust AI to catch counterfeits | The Federal Reserve verifies deposited cash “note-by-note, on sophisticated processing equipment”, and people only judge the notes the machines reject | Protocol 1 |
| A 100 millisecond speed improvement increased conversion rates significantly | A Google-commissioned Deloitte study of 37 brands found a 0.1 second improvement was associated with 8.4% more retail conversions and 10.1% more travel conversions, and it’s observational and from 2019 | Frontend already has external standards |
| Dieter Rams advises people to write down what makes good things good | All I can find is Rams writing his own standard down: 3 rules in 1975, 6 principles in 1983 and 10 by 1985 | Protocol 4 |
| My MCP App also returned text, so it’s accessible | The spec requires that text for the model and for text-only hosts, and the interface you render still has to be keyboard and screen reader accessible | What are MCP Apps? |
What I still don’t know
At the end of a different talk, How to Thrive as a Professional with AI at React fwdays 2025 (recording), I raised a question I couldn’t answer. That talk is about invariants, the things that stay constant while the tools change, and I borrowed Noether’s theorem from physics for it, loosely: when you see a symmetry from a variety of angles, there’s an invariant there. A lot of people say taste is the thing that separates humans from the machines, so is taste invariant? If you look at taste from a wide variety of angles, is it constant, and is there some type of law around taste? Or is taste just the aggregation of the majority? I said in that talk that I honestly had nothing to share there.
Cutting’s study says part of what we call good taste really is the majority’s exposure, and the expertise studies say trained people see things the rest of us don’t, and I don’t know how to weigh those against each other.
As far as I can find, nobody has run the experiment that would settle the practical question either: take developers, train half of them for a few months with something like the protocols above, and have designers rate everyone’s work blind. If you have thoughts on any of this, please reach out.
Takeaways
- Taste is a trainable skill for tacitly judging quality against external standards. The external part is what separates it from preference and from bias.
- Don’t trust a gut call just because it feels sure. Trust it when it was trained somewhere with reliable cues and clear feedback, and check it when it wasn’t.
- Pair every example with an outcome. Immersion alone can teach you to like what’s common.
- Write your standard down with examples, anchor it to Core Web Vitals and WCAG 2.2, and hand the same file to your agent as a taste skill. The file only gives the agent the standards in writing and you still have to judge what comes out.
- Treat the standards as the floor. Failing them usually makes something distasteful, and a page can pass all of them and still be tasteless.
I’m not here to prescribe things: these are ingredients for a dish that you’ll cook. Train the judgment, write it down, and check it against something outside yourself.
I gave this as a talk at Future Frontend 2026 in Espoo, Finland, and the recording has the live demos:
If you’re programming an event and would rather have this delivered than read, here’s what I speak about and how to book me. If you’d rather have your team taught than your audience addressed, I run workshops. If any of this was interesting or useful, please share it with someone who might benefit.
ok bye
Questions
What is good taste in software engineering?
Good taste in software engineering is a trainable skill: the ability to tacitly judge the quality of software against external standards. You can tell an interface or a piece of code is off before you can say why, and the judgment is anchored to something outside your own preferences, such as Core Web Vitals, WCAG 2.2 or a written set of principles.
How do you develop good taste?
You develop good taste with practice that has fast, clear feedback attached. Immerse yourself in examples that come with a measured verdict, compare before you read the standard, take corrective feedback, write your standard down, borrow from other fields and look at everything around the thing you built. Kahneman and Klein's 2009 paper says gut judgment becomes trustworthy only where cues reliably predict outcomes and only with enough practice.
Can taste be taught?
Yes, taste can be taught, within limits. A 1993 study of art expertise found training strongly associated with better aesthetic judgment, and in a 1987 study researchers wrote down what 1 veteran chick sexer (someone who sorts day-old chicks into male and female) looked for and novices' agreement with the experts went from a correlation of .21 to .82. It takes a while though: a 30 minute training video produced only a small effect in a 2015 study.
Is good taste innate or learned?
The research says good taste can be trained, and that talent still matters. A 1993 study of 5 levels of art expertise found training strongly associated with better aesthetic judgment, and Kahneman and Klein's 2009 paper says skilled intuition needs an environment where cues reliably predict outcomes plus enough practice in it. The same paper says talent surely matters, so the research doesn't promise that everyone gets to great taste.
Is taste subjective or objective?
Preference is subjective, and taste is judgment against standards that exist outside the person judging. For art in general people still argue about whether such a standard exists, but for frontend a lot of it is published and measurable in Core Web Vitals and WCAG 2.2. Those standards are the floor, since a page can pass both and still be tasteless.
How do you know if you have good taste?
You can't tell whether you have good taste from how confident you feel. Kahneman and Klein's 2009 paper says there's no subjective marker that separates a correct intuition from one produced by a bad heuristic, so the test is whether your judgment was trained somewhere with reliable cues and fast, clear feedback. A practical check is to commit to a call before you see a measured verdict, such as a Core Web Vitals or WCAG result, and keep a tally.
Who said taste is the new core skill?
Greg Brockman, OpenAI's president, posted "taste is a new core skill" on February 16, 2026. He wrote "a", and people now search for it as "taste is the new core skill". Paul Graham had posted a similar prediction 2 days earlier, on February 14, 2026, and the Linux kernel maintainer Jens Axboe replied that taste has always been a core skill and is just easier to see now.
What is taste in the age of AI?
Taste in the age of AI is the same skill, and it's easier to see because generating got cheap. Sundar Pichai said in April 2026 that 75% of new code at Google was AI-generated and approved by engineers, so the reviewing and approving is still people. In Stack Overflow's 2025 developer survey 66% of developers said they're frustrated by AI solutions that are almost right, but not quite.
Can AI have taste?
AI can't have taste by itself, as far as the evidence goes. Without direction, a model samples the most common choices in its training data, which Anthropic calls distributional convergence, and a taste skill just gives a coding agent the standards in writing. When the SkillsBench authors had agents write their own skills, the self-written ones scored 8.1 to 11.5 points below using no skills at all, so somebody with trained judgment still has to write the standard and judge output against it.
What is a taste skill?
A taste skill is a markdown file of written design standards that an AI coding agent loads before it builds UI, so that its output can be judged against them. It only gives the agent the standards in writing, so a person still has to judge what comes out.
What is tacit knowledge?
Tacit knowledge is what you know and can't fully put into words. The term is Michael Polanyi's, from his line that we can know more than we can tell. A lot of it can still be written down: researchers interviewed neonatal nurses incident by incident and turned their cues for sepsis into training material.
What are MCP Apps?
MCP Apps is the first official extension to MCP (the Model Context Protocol): an MCP server ships its own interactive interface, and the host (an AI app like ChatGPT, Claude or VS Code) renders it in a sandbox inside the conversation. It went live on January 26, 2026. The rendered view still has to be keyboard and screen reader accessible, and the spec doesn't do that for you.
Written by me, Tejas Kumar, an AI Engineer at IBM based in Berlin. Read everything else I have written, or go to Fluent React, my O'Reilly book on how React works inside, the talks I give at conferences, and ConTejas Code, my podcast.