AI Deployment
Feature 20  ·  Quality  ·  Edition Q1 2026

Designers cannot agree
whether design got better.

Asked whether design has improved, the profession splits almost exactly in three. That is not a mixed result. It is the absence of a shared standard - and it explains why the other five features on this desk describe problems nobody can currently settle.

One question, asked of designers: has design got better?

36% said better. 35% said worse. 29% said the same.1

Figure 01
Has design got better?
Designers, on the direction of their own profession's output.
Single stacked bar showing designers split 36 per cent saying design has got better, 35 per cent saying worse and 29 per cent saying the same.
Source: design industry research, 2026. The three figures sum to 100, so this is a partition of one question rather than separate findings. We do not have the question wording, which matters - "design" could be read as the discipline, the output, or the working conditions, and the three would produce different answers.

A one-point gap between "better" and "worse" is not a verdict. On any reasonable sample it is a tie, and a tie on this question is a more interesting result than either answer would have been.

Ask engineers whether software got faster and you get a number. Ask designers whether design got better and you get a coin toss. The difference is not temperament. It is that one profession has an agreed measure and the other does not.

What a tie actually tells you

There are three ways to read a 36/35/29 split, and they are not equally likely.

The first is that design genuinely got better for some people and worse for others - different sectors, different employers, different kinds of work. Plausible, and probably partly true.

The second is that respondents answered different questions. Some heard "is the craft better", some heard "is my job better", some heard "is the output in market better". The absence of published question wording means we cannot exclude this, and we flag it in Figure 01 rather than assume it away.

The third reading is the structural one. The profession has no agreed instrument for answering the question at all, so what the survey captured was not an assessment but a mood - and moods distribute roughly evenly in a population under pressure.

We think the third is the most useful, because it is testable against everything else on this desk.

The absence explains the rest of the edition

Five features precede this one. Each describes a problem that a shared quality standard would resolve, and that without one cannot be settled.

Figure 02
What each desk problem needs in order to be answerable
The Creative desk, Edition Q1 2026.
Feature 13 - the average Cannes winner scored 2.2 with consumers
Expert juries and consumer testing produced different orderings of the same work. Neither is wrong; they measure different things and no standard reconciles them.Needs: an agreed definition of "good"
Feature 16 - you can make a hundred and read thirteen
Eighty-seven variants get killed on judgement rather than evidence. Whose judgement, applying what criteria, is undefined.Needs: a defensible basis for selection without data
Feature 17 - Adobe grew 12.7% and fell 60%
The market is pricing whether professional design capability remains scarce. A profession that cannot state what its expertise produces cannot argue the point.Needs: an articulable claim about what craft adds
Feature 18 - the job spec was rewritten first
Listings specify tools because tools can be named. The scarce capabilities - brief-writing, editorial judgement - resist specification.Needs: a way to describe judgement in a job posting
Feature 15 and 19 - accessibility and disclosure
Two regimes now impose external, testable standards on design output - contrast ratios and marking obligations. They are the only agreed criteria in this edition, and they arrived from outside the profession.Needs: nothing. They are already specified
This mapping is ours. It is an argument about what the desk's findings have in common, not a result any source reports. The last row is the uncomfortable one and we state it deliberately.

Read that last row again. The only enforceable quality criteria arriving in design this year came from regulators. A contrast ratio of 4.5:1 is a standard. A marking obligation is a standard. They are narrow, they are external, and the profession did not write them.

Why this profession in particular

It would be easy to read the split as ordinary professional grumbling. It is worth asking why design specifically lacks the instrument that neighbouring disciplines have.

Performance marketing has cost per acquisition. Engineering has latency, uptime and defect rates. Copy has readability scores, however crude. These are not perfect measures and practitioners complain about all of them - but a complaint about a measure is a different condition from having none.

Design's difficulty is structural rather than cultural. Its output is judged on effects that are separated from it in time, mixed with other causes, and often not measured at all. A layout contributes to a conversion rate alongside the offer, the price, the traffic source and the season. Isolating its contribution requires exactly the kind of controlled comparison that Feature 16 shows most organisations cannot run past thirteen variants.

There is a second difficulty and it is the more interesting one. Much of what design does is avoid outcomes rather than produce them. A clear interface prevents confusion. A well-set page prevents abandonment. A correct contrast ratio prevents exclusion. Prevention leaves no trace in the data - the support ticket that was never raised does not appear in any system - and a discipline whose value is largely in absences will always struggle to demonstrate it.

Design is measured on what it produces and paid for what it prevents. The gap between those two is where the argument about quality lives.

The consequence inside an organisation

An unmeasured discipline does not become unmanaged. It becomes managed by proxy, and the proxies are worse than the thing they replace.

In the absence of a quality standard, creative decisions get resolved by seniority, by volume of stakeholder opinion, by whoever briefed the work, or by what resembles a competitor. None of these is a judgement about the work. All of them are stable, repeatable and defensible in a meeting, which is precisely why they persist.

This is also the mechanism by which the effectiveness gap in Feature 13 becomes possible. Award juries and consumer testing produced different orderings of the same films - but an organisation with no standard of its own has no basis to prefer either, and will default to whichever is more socially useful. Awards are more socially useful.

What the profession is being asked to do instead

The sentiment research describes designers as taking on a wider remit - moving into ambiguity rather than away from it.1 Read alongside Feature 18, that expansion has a specific shape.

Feature 18 found the scarce capabilities to be brief-writing, editorial judgement at volume and constraint design - none of which appear in job specifications, all of which resist description. Those are the capabilities of someone deciding what should exist, not someone executing a decision already made.

That is a promotion in substance. It is also a promotion into exactly the territory where the missing instrument bites hardest. A discipline moving upstream into judgement, without an agreed way to evaluate judgement, is taking on more responsibility and less defensibility at the same time.

We do not have data on how that resolves. Nobody does - the transition is roughly two years old. But it makes the case for a written house standard more urgent rather than less, because the alternative in an ambiguous remit is not neutrality. It is seniority.

The sentiment data does not resolve it

Designers are not, in general, gloomy about the tooling. 67% view AI as a complement rather than a replacement. 91% say AI tools improve their designs.1

Hold that against the split in Figure 01 and something does not sit.

On their own tools
91%
Say AI tools improve their designs. Near-unanimous, and about work they can see directly.
On the profession's output
36%
Say design has got better. A near-tie, and about work in aggregate.

Ninety-one per cent report improvement in the thing in front of them. Thirty-six per cent report improvement in the thing overall. That is not a contradiction - it is what you would expect when individuals can assess their own work and nobody can assess the whole.

It also has a less comfortable reading, and we cannot distinguish between them: it may be what a population looks like when everyone believes their own output improved and is looking at everyone else's.

Whose research this is, and ours

This sentiment data is published largely by Figma, which sells design tooling and has an interest in designers reporting that new tools improve their work. We have used the figures because they are the only systematic sentiment research in the category, and marked the interest rather than absorbed it. Our own position is worse, not better: Marketing Legendary sells creative work and this feature argues that judgement is scarce and unmeasured. That is a convenient thing for us to argue. The reader should discount accordingly - which is, in miniature, exactly the problem the feature describes.

What would actually settle it

We are not going to propose an industry standard for design quality. Attempts at that have a poor record and we have no evidence any of them worked.

What can be said is narrower and more useful: the question is answerable at the level of a single organisation, and almost nobody answers it.

Figure 03
Two ways to hold a quality standard
What is available to an individual organisation.
Not available
And unlikely to become available
An industry definition of good designRepeatedly attempted, never settled. The 36/35/29 split is what its absence looks like.
Award outcomes as a proxyFeature 13 established that jury ordering and consumer response are different orderings.
Aggregate profession-level measurementNobody is positioned to conduct it and no instrument exists.
Available now
To any organisation willing to write it down
A stated house standardWhat this organisation believes good work does, written down, applied in review, and argued about openly when it fails.
Outcome measurement within your own ceilingFeature 16's calculation gives you the number of variants you can genuinely evaluate. Below that ceiling, evidence is available.
External specified criteriaContrast, marking, performance budgets. Narrow, but real, and already enforceable.
This framing is ours. No source proposes it. It is an argument that the standard-setting problem is unsolvable at industry level and tractable at organisation level, and we offer it as an argument rather than a finding.

A house standard is not a definition of good design. It is a statement of what this organisation is trying to make happen, specific enough that a piece of work can fail it. Most creative organisations do not have one, which is why creative review so often resolves into whoever is most senior in the room.

What to do about it

Write down what your organisation thinks good work does. Not adjectives - effects. What should a viewer understand, feel or do. A standard that cannot be failed is not a standard.

Separate craft assessment from outcome assessment, and run both. Feature 13 found expert judgement and consumer response producing different orderings. That is information, not a problem, provided you stop treating either as a proxy for the other.

Use the external standards you have been given. Contrast ratios and marking obligations are the only agreed criteria in this edition. They are narrow, but they are testable, and a team that meets them reliably has more shared standard than most.

Measure inside your ceiling and accept judgement above it. Feature 16's arithmetic tells you where evidence stops. Beyond that point, decisions are made on taste - and a team that names that explicitly makes better decisions than one that pretends the data continues.

Stop citing the split as bad news. A profession divided on whether its output improved is a profession without an instrument. That is a solvable problem locally and it does not require anyone to agree with anyone else.

Figure 04
Design quality, Q1 2026
What designers report about their tools and their profession.
MeasureValueNote
The profession's own verdict
Design has got better36%Reported
Design has got worse35%Reported
Design is the same29%Reported
Gap between better and worse1ptOur arithmetic
On their own tools
Say AI tools improve their designs91%Self-reported
View AI as complement, not replacement67%Self-reported
The only external standards arriving
WCAG 2.1 AA contrast and operabilityEnforceable since 28 June 2025 - Feature 15
AI content markingBinds 2 August 2026 - Feature 19
Sentiment figures are self-reported and drawn from vendor-published research. The two external standards are legal instruments, not professional ones.

How we did this

Where this comes from
A named study, reported by someone else: design industry sentiment research, 2026, for the quality split and tool sentiment. No Tier 1: we did not obtain the survey instrument, sample size, sampling frame or question wording for any figure.
Our own arithmetic
The one-point gap is arithmetic. The reading that a tie indicates an absent instrument rather than a mixed outcome is our argument, offered with two competing readings stated alongside it.
What's ours, not the source's
Figure 02's mapping of desk problems to the missing standard, and Figure 03's industry-versus-organisation framing, are ours. No source proposes either.
Interest
Disclosed in the body, on both sides and unusually directly. The research is vendor-published; our conclusion flatters our own business.

What this doesn't prove

  • That design has or has not improved. The feature argues the profession cannot answer this. It does not answer it either.
  • That respondents interpreted the question consistently. Without published wording we cannot exclude that the split reflects three different questions being answered.
  • Sample size or margin of error. Unavailable to us. A 36/35 gap may be well inside sampling noise, which would strengthen rather than weaken the reading offered here.
  • That a shared standard would improve outcomes. Plausible and untested. Professions with agreed measures also optimise toward them, sometimes badly.
  • That house standards work. We recommend them on reasoning, not evidence. No research we found tests whether organisations with written creative standards produce better work.
  • That design is harder to measure than neighbouring disciplines. The prevention argument is reasoning about the nature of the work, not a comparative study of measurability across professions.
  • That proxies like seniority actually dominate creative decisions. This is our characterisation of how unmeasured disciplines behave. We found no research observing creative decision-making in organisations.
  • That the 91% and 36% figures come from the same respondents. They appear in the same body of research but we cannot confirm they are the same sample, and the comparison in the twin panel assumes they are.

Sources for this feature

  1. Design industry sentiment research, 2026. figma.com, figma.com A named study, reported by someone else - vendor-published, interested party
  2. Features 13, 15, 16, 17, 18 and 19 of this edition. Another feature in this edition
End of the Creative desk
Q1 2026 · Six Features
Marketing Legendary publishes quarterly. The next edition follows.
DA
The practice behind this desk

designs.art

We write the standard down before the work starts, so that a review has something to fail against other than the most senior opinion in the room.