Designing with AI

With AI agents, you now have to think like three professionals

A digital product used to move through product, design and engineering, one after another. Today agents do much of the work of all three, but the judgment of each craft is still yours. Here's what I learned, with a real example and tools you can use tomorrow.

Thiago Soares · 8 min read

20202026

In 2020, a page like my portfolio homepage went through three people in a row. Someone in product decided what to say and to whom. Someone in design decided how it would be understood. Someone in engineering built it and made sure it didn't break. Each one owned a part, and the next person caught the previous one's mistakes.

This week I redesigned that homepage with AI agents. Claude Code designed and wrote the code; Codex reviewed it. It moved fast, and along the way I learned something I didn't expect: agents do a lot of the work of all three crafts, but they don't bring the judgment of any of them. If you don't bring it, nobody does.

Product

In 2020
One person decided what to say and to whom.
Today the agent
Keeps producing versions, even when the request is wrong.
Still your job
Deciding what must land in 3 seconds.

Design

In 2020
Another decided how it would be understood.
Today the agent
Takes adjectives literally and swings to extremes.
Still your job
Format, word limit and who it is for.

Engineering

In 2020
Another built it and made sure it held up.
Today the agent
Writes the code and calls it done.
Still your job
Having another model review it, and asking for proof.

What follows are four lessons from that project, grouped by craft. Engineering gets two, because that's where I made the most mistakes: reviewing and testing. Each has what happened, a piece you can play with, what I'd do differently and a template to copy. At the end you can build your own request using the three roles.

Product

1. The hard part isn't making. It's deciding

The third animation on the homepage went through eight versions in one morning. One with lots of text. One made of abstract dots. One with icons. Another with speed bars. They all moved nicely, and none of them explained what I wanted to show.

The agent wasn't the problem. It did exactly what I asked, and what I asked was badly framed. In 2020, someone in product would have asked the question before anything got drawn: what does a visitor need to understand in three seconds? After the eighth version I stopped and wrote that sentence down. I picked three seconds myself, thinking of an audience that skims. The next version used a Gantt chart to show how I went from working in a line to working in parallel, and it finally said what I'd been trying to explain.

Illustration
Eight versions, one question. Only the one that passes the 3-second test stays.

An agent rarely asks why unless you tell it to, which is why the template below does. And when a new version costs a minute, it's tempting to ask for another one instead of figuring out what went wrong.

What I'd do differently

  1. Before asking for anything, write in one sentence what the person needs to understand.
  2. Start with three options. If none of them gets the idea across, fix the request before asking for more.
  3. Judge each option against that sentence, not against how I feel that minute.
Template
Context: [what it is and who it's for].Whoever sees this must understand in 3 seconds that [the sentence].Before proposing anything, tell me whether that sentence is clear or what it's missing.Then propose 3 different options, say in one line whether each meets the sentence, and recommend one.

Design

2. AI agents go to extremes

It happened to me again and again with the animations. When I asked for something "simpler", the agent stripped so much that the drawing stopped meaning anything. When I asked for "clearer", it filled the screen with text. It swung from one end to the other and never landed in the middle.

What was missing was design judgment: adjectives aren't instructions. What helped was swapping them for concrete decisions. I picked formats people recognize without explanation, like a dashboard, a counter or a screen with a cursor. I capped the text at three words per animation. And I told the agent who it was for: someone who comes to the site to evaluate my work and won't stop to read.

abstractall text

Too abstract: nobody gets what it is

Drag the slider. Agents tend to land at the ends.

What I'd do differently

  1. Don't ask for "simpler" or "clearer". Ask for a specific format: dashboard, list, counter, before and after.
  2. Put a word limit in writing.
  3. When the answer swings to one extreme, name the other one as the limit: "simpler, without making it abstract".
Template
Show [the idea] using a familiar format: [dashboard / counter / list / before and after].[number] words max inside the graphic.Audience: [who]. They need to get it without reading a paragraph.Don't make it abstract and don't fill it with text.

Engineering

3. Whoever writes it doesn't review it

This is the biggest change in how I work. I use two agents from different companies: Claude Code, from Anthropic, writes the code and the copy; Codex, from OpenAI, reviews it with one instruction: try to reject it. Nothing ships without its approval.

In my experience, two different models get different things wrong. The homepage chat had two bugs that Claude Code had called done. If someone left the page just as they unlocked a secret, the chat got stuck on "typing". If the connection dropped at the wrong moment, the retry button disappeared. Codex found both.

It wasn't the only time. For this article alone, Codex flagged 25 Spanish sentences that sounded AI-written. In the code review it asked for six changes; five were real, and it only approved once they were fixed. That's the job a demanding engineering teammate used to do.

The code looked fine to the agent that wrote it.

What I'd do differently

  1. Always review with a different model from the one that wrote the work, and give it the cases I want checked.
  2. Ask it to look for problems, not to say whether it's fine.
  3. If it flags a possible bug, ask how to reproduce it. If there's no way, ask what it saw.
Template
git diff | codex exec "Review this change as if you were going to reject it. You didn't write it.Look for: edge cases, what happens if the user leaves halfway, what happens if the connection fails, copy that doesn't make sense.Reply with a list ordered by severity, with file and line. End with APPROVE or CHANGES-REQUESTED."

4. If I didn't test it, it isn't done

The other half of engineering is proof. Before publishing, an automated browser checked the homepage at desktop and phone sizes, in the site's three languages, and returned screenshots and measurements of the main pieces.

Those measurements caught two mistakes. I'd been told a change was halfway done, and measuring showed it had never been started. Another time, on the phone, a fixed control showed up on top of the chat. Neither one appeared in the agent's summary. The screenshots let me compare what the agent said with what was actually on screen.

Desktop · ES
Desktop · EN
Desktop · PT
Phone · ES
Phone · EN
Phone · PT
Desktop and phone, in the three languages of the site. Measured, not assumed.

What I'd do differently

  1. Write down what "done" means before starting: which screens, which languages, which cases.
  2. Ask for proof you can look at: screenshots, measurements, logs.
  3. Look at those screenshots myself, not at the summary.
Template
Before saying it's ready, test on desktop (1440 px) and phone (375 px), in [languages].For each combination, give me a screenshot and one line on what you measured.If something fails, don't call it done: tell me what failed and what you'll change.

What didn't work

With the sound, I was missing design judgment, and I didn't see it in time. I asked an agent to add sound effects to a short video. It handed back the file with volume and sync measurements and called it done. When I played it, the synthetic beeps sounded cheap.

I had accepted the measurements before listening. For the next version I used music and sound effects made by people, with each effect timed to a movement on screen. Now I listen to, read and look at everything that depends on taste myself. No measurement tells me whether it sounds like me.

Build your request with the three roles

This request gathers the decisions of all three roles before you start. Fill in the fields and paste the result into Claude, Codex or whatever AI you use. Then run what you get back through the three reviews below.

Product
Design
Engineering
Your request, ready to paste into any AI
Fill in the fields and press "Build my request".

Review it as product

Template
You're a product lead. Read this and tell me in 3 lines: who it's for, what that person understands in 3 seconds, and what they still need in order to decide. If the answer isn't clear, say so.

Review it as design

Template
You're a senior designer. Check whether this can be understood without reading: format, hierarchy and amount of text. Flag what's extra and what's too abstract. Propose one change, the most important one.

Review it as engineering

Template
You're an engineer and your job is to break this. List what can fail: edge cases, phone, slow connection, languages. Order by severity and explain how to reproduce each problem.

What I'm taking with me

Agents don't replace product, design or engineering. They do much of the work, but each craft's judgment has to live in someone. On this project, that someone was me, with a second model reviewing everything.

That's why understanding all three crafts matters to me today as much as mastering my own. A designer who knows what product would ask and what engineering would break works better with agents, and with people too.

I'm still learning to work this way. If you work with agents too, I'd like to know how you check what they give you.

Want to see how I applied this?

Nova, the AI on my homepage, knows the project. Ask it anything, or look at the cases.

See my work
← Back to all articles