Directing AI agents to build my portfolio, and the failures that taught me to run it

I wanted a portfolio that worked like a product, not a page. I am a designer, not an engineer, so I ran the build the way I run a design team: clear briefs, short sprints, and a hard review of everything that came back. Claude Code was the build team. The first draft took 3 to 4 hours. The full first version went live in 7 days across 13 sprints and 61 commits, with a custom CMS, a guarded AI agent, SEO and AI-search metadata, and its own analytics. Then things started to break. Free services paused, AI models were retired, a token expired, and the agent went silent without telling anyone. Building with AI turned out to be the cheap part. Deciding what to build, reviewing what came back and keeping it running were the work.

Year
2026
Type
Product Leadership / AI Adoption
Domain
Personal platform (portfolio, CMS, AI agent)
Role
Product owner, design director and reviewer. AI coding agents did the build.
Team
One person, plus Claude Code as the build team
Timeline
First draft in 3 to 4 hours. Full first version in 7 days (11 to 17 March 2026). Hardened in June. Rebuilt key parts in September and October. Current release: v2.9.0.
Directing AI agents to build my portfolio, and the failures that taught me to run it

1 October, the day the agent went quiet

A visitor asked the chat agent about my work. It apologised and gave nothing back.

The logs showed all three AI services failing at once. One account had run out of credit. The other two services had retired the models the code asked for. The site had no way to notice, so for a period I cannot measure, every visitor got an apology instead of an answer.

That outage is the best summary of this project. Nothing about it was hard to build. All of it was hard to run. The rest of this page is the logbook of how the platform got here, one failure and one rule at a time.

What I set out to build

Most portfolios are screenshots with captions. I wanted one a hiring manager could question, and get a straight answer that does not leak anything private.

Two constraints shaped everything: no engineering background and no engineering budget. Whatever I built, I had to own, change and fix alone. So the brief was:

  • Content I can edit without touching code. A CMS behind a login.
  • An agent that knows my work and nothing else. Public facts only, with rules it cannot be talked out of.
  • Findable by people and by AI search. Proper metadata, a sitemap and an llms.txt file for AI crawlers.
  • Cheap to run. Free tiers where possible.

Who did what

  • I owned: the brief, scope and priorities, the information architecture and content, every design decision, review and acceptance of every change, and what to cut.
  • I directed: all code, written by Claude Code in the terminal from my briefs, in short sprints.
  • I did myself where the tools could not: signing in to services, approving access, entering passwords and keys. AI agents should not hold credentials, so I kept that boundary.
  • I did not: write the code by hand. I read it through the AI's explanations, not line by line.

I worked the way I lead a team. A goal and a constraint, not a solution. I reviewed results on screen, often by selecting the exact element that was wrong and saying why. Anything weaker than the brief went back.

March, week one: too fast, too much

What happened. I chose plain HTML, CSS and JavaScript with no build step, plus a small set of server functions for the parts that need secrets. The fashionable choice was a framework. Mine was faster to execute. The first commit describes a 46KB home page, a 54KB admin page, accessible to WCAG AA, and "zero build step". The first version went out on GitHub Pages, which only serves files. The chat agent needed somewhere to keep its API key, so the same day the site moved to Vercel, where small server functions can hold secrets.

The week ran as 13 numbered sprints and 61 commits, each with one theme: the CMS, a visual case study editor, analytics, SEO and AI-search optimisation, a motion system, a theme editor. Then the excess. The background music went through five versions in three days: a generated ambient soundtrack that made no sound, a Spotify embed, a version that reacted to scrolling and clicks, a "dark soul R&B" generator, and finally one MP3 behind a play button. The player shrank from 457 lines to 92. Particle effects grew to five presets. Sound effects went from off by default to always on.

The rule it left. Building with AI is so cheap that the brake has to come from judgment, not effort. The question moves from "can we build it" to "should we".

March, day one and day seven: secrets kept leaking into the page

What happened. The first chat agent called the AI service straight from the browser, which would have exposed the API key to anyone who looked. Within hours it went through a server function. For the first six days, the CMS password sat inside the admin page's own code. On 17 March it moved to the server, which now checks the password and issues a login token that expires daily. Every admin function checks that token.

The rule it left. Nothing secret lives in the page.

14 and 16 March: features that looked finished

What happened. The CMS login stopped working. Three overlapping versions of the same function had piled up across separate editing sessions and clashed. Two days later, images I had "uploaded" in the CMS turned out never to have reached the live site. Around the same time, publishing large images hit the hosting platform's 4.5MB request limit, and the fix split uploads into smaller batches.

The rule it left. From then on, the sprint notes record a syntax check of the scripts before they ship. Looking done is not done.

June: the database fell asleep

What happened. The database service paused itself after a quiet period, and every deploy failed. Analytics moved to plain file storage that cannot pause.

The rule it left. Find the dependency that broke, remove it from the critical path, and leave a fallback behind.

September: two limits and an expired token

What happened. Image uploads failed because the hosting platform caps upload size, the same 4.5MB limit from March. This time images stopped going through that path at all. They go straight to a storage bucket on another provider, through a small gatekeeper service that checks my CMS login. Then publishing broke because a deployment token had expired. Creating a new one needs a human, and I could not do it that night. Rather than wait, publishing moved to the same storage bucket, so a case study goes live when I press Save.

When the bucket and gatekeeper had to be set up, I said plainly that I did not know how. The AI drove my browser and desktop apps to do it, with my approval at each step. Its own safety checks stopped it twice: creating server code from my browser, and opening the page that issues GitHub tokens. One went ahead after I gave explicit approval. The other we routed around.

The rule it left. Do not wait on a fix that needs a person who is not there. Move the work.

1 October: never silent, instead of never failing

What happened. The outage at the top of this page. The backup plan dated from 13 March: if one AI service failed, try a second, then a third. Its flaw was that each service was pinned to one model name, and model names get retired. Three providers, one point of failure each.

My first instinct was to demand that it never fail again. Nobody can promise that, because the AI services themselves go down. So the goal changed:

  • Each service now has backup models. If one is retired, the agent tries the next, and if the whole list is stale, it asks the service which models it offers today.
  • Every call has a time limit, so a slow service hands over instead of freezing.
  • If every AI service fails, the agent still answers from the published knowledge file, without AI. A visitor asking for my resume gets the link, and is told plainly that the AI is offline.
  • Failures are logged with the real reason, which is how the out-of-credit problem was found in minutes.

The rule it left. "Never silent" is a promise you can keep. "Never fails" is not.

October: the guardrails were editable by visitors

What happened. The rules that keep the agent private were sent from the web page to the server. Anyone who edited the page in their browser could have removed them. They moved to the server, where a visitor cannot reach them. The page can add context. It cannot switch the guardrails off.

The agent now refuses private details such as phone numbers or family information, declines unrelated tasks, and ignores attempts to rewrite its instructions. I tested it on the live site with a question asking for my home address. It declined and gave my city, which is public.

The rule it left. Rules live where the user cannot touch them.

October: two widths asking for two different things

What happened. I reviewed the site on a laptop and a phone side by side. The navigation squeezed between 960 and 1,280 pixels, the logo overlapped the menu, and the resume button broke onto two lines. I set the rule: below 1,280 pixels the menu collapses into an animated menu button. Then the details the AI had missed. The menu button's lines were heavier than the theme button's, the boxes were different sizes, and the menu text had taken a bold, all-caps style from the source component. Each was measured and fixed: matching 40-pixel boxes, matching line weight, the site's own typography.

The two animated menu components were written for React. Instead of rebuilding the site around them, the AI rebuilt the same motion in the site's own code. The plain-pages choice from March paid off.

Then the bigger catch. The phone menu ended with "Contact". The desktop menu still ended with "View resume". I pointed it out, and the AI admitted it had left the desktop alone on purpose. "Contact" is now the single call to action at every width, and the hero carries two actions instead of three.

The rule it left. One call to action, everywhere.

October: recruiters brought their own context

What happened. Recruiters paste a full job description or send a link. The agent cut messages at a fixed length and could not open a page. Now a visitor can paste up to 6,000 characters, with a counter once they pass half of that, and up to two links per message, which the agent reads before it answers.

Reading links opened a new risk. A page can carry hidden instructions, and a link can point at the server's own private addresses. So the rules were set before the feature shipped. The server only fetches public web addresses, re-checks every redirect, and stops at a size limit. Page text reaches the agent as data, clearly fenced, with an instruction never to follow anything written inside it. I tested it on a real job posting. The agent read the role and walked through the fit.

The rule it left. Decide the guardrails before the feature, not after.

October: the agent may not learn from strangers

What happened. Once the agent could read anything a visitor sent, a bigger question followed: should it learn from those conversations? No. Nothing a visitor says can change what the agent knows about me. If someone claims a fact about me, it says it cannot confirm it.

I still wanted to know what happened. So the agent raises an alert in the CMS when:

  • a visitor, or a linked page, tries to override its rules
  • someone asks for private details, such as salary or a home address
  • a reply starts to echo its internal instructions (the reply is replaced before the visitor sees it)
  • it is asked something about me it cannot answer
  • every AI service fails and it falls back to a basic answer
  • a shared link cannot be read

Each alert comes with a next step, and the CMS builds a plan from the open ones: tighten guardrails first, then fill the most-asked gaps. I tested it with an "ignore your instructions" message and a salary question. Both arrived as alerts, and the agent declined both.

Storage is split in two, on purpose.

What it holds How long Who can add
Learning Facts I approve, and a monthly history of alerts and unanswered questions Permanent Only me, in the CMS
Logs Raw alerts and visitor visits Deleted after 30 days The system

When a gap alert comes in, I type the answer and press one button. The agent uses it within two minutes, without a deploy.

The first live test of the alerts recorded nothing. The storage service holding them had been suspended, most likely because visitor analytics wrote to it every 60 seconds from every open tab. Failed writes were swallowed, so nothing said so. Analytics had stopped saving too. The same silent failure, one level down. Alerts and visits moved to the bucket that already held the case studies, and analytics now writes once per real visit, skipping visits under 5 seconds and automated browsers. When the AI tried to update the gatekeeper service from my browser, its own safety check blocked it as a production deploy. It stopped and asked. I approved it in my next message.

The rule it left. Let users inform the system, never rewrite it.

October: optimistic, not desperate

What happened. While writing the agent's hiring guidance, the AI found an old line in the knowledge file saying roles needing relocation outside Southeast Asia were outside my interest. I am open to the right role abroad. The agent would have turned away the conversations I wanted most.

The line went. In its place, a stance: connect my strengths to what the visitor needs, with evidence. Never rule out a senior role over location, pay or a partial match. Never plead, discount, quote rates, say I am job hunting, or promise I will accept. I tested it with a Head of Design role in Berlin, with relocation and no fixed budget. The agent called it worth a conversation and pointed to a call. It also offered to "flag it to him directly". It cannot contact me, so that promise was false. A rule now stops it.

The rule it left. An agent that speaks for you needs a stance, not only facts.

October: the version number froze in March

What happened. The CMS still said 1.1.1 after more than forty changes. I noticed, not the AI. Two bugs sat behind it: only deploys from the CMS raised the number, and the CMS never read the version file. The AI rebuilt the history from the commit log into Semantic Versioning: 32 releases from v1.0.0 to v2.6.4, each tagged on GitHub. The single major release, v2.0.0, marks the day publishing moved off GitHub.

The rule it left. A version that only moves when one tool deploys is not a version.

4 October: case studies become pages

What happened. Case studies opened in an overlay on the home page, which made them hard to share and invisible to search. Each one now has its own address and page, and I can drag them into the order I want in the CMS.

The rule it left. If a recruiter cannot send it to a colleague in one link, it does not exist.

The scoreboard

  • A first draft in an afternoon, a platform in a week. 3 to 4 hours to a first draft. 7 days, 13 sprints and 61 commits to the full first version.
  • A live platform I own end to end. A portfolio, a private CMS, case studies published from the CMS, a guarded AI agent that reads job links and flags what it cannot answer, SEO and AI-search metadata, and analytics. About 10,300 lines of code across the core files, none typed by me.
  • A release history I can point to. Every release since v1.0.0 in March is tagged, with a one-line summary.
  • Readability fixed where I found it lacking. The chat panel's footer text went from a contrast ratio of 1.6:1 to 5.4:1, which meets WCAG AA.
  • The method travelled. I used the same approach to build two more sites with AI agents: a photography portfolio in June and another web product in July, each with its own CMS.
  • It changed how others work. I ran an internal learning session at work on building a CMS with AI. A colleague went on to build and ship a side project of their own.
  • Still open. A deployment token needs regenerating, and one AI account needs credit. Both need me, not the AI.

Rules I run by now

  • Monitoring on day one. Alerts record outages, but only when a visitor triggers one. A daily test question would catch it first. I built resilience after an outage. It should have come before launch.
  • Version from the first commit. It would have cost nothing.
  • A smaller first week. Half the sprint themes would have produced a better site. The music, the sound effects and the extra integrations cost review time I later spent removing them.
  • Write the design checklist down. Matching icon sizes, one call to action and consistent typography are rules I held in my head. Written down, the AI would have checked them before I had to.
  • Brief the goal, not the solution. The best results came from a constraint and a reason.
  • Review on the real surface. Selecting the broken element on screen and saying why beat any written ticket.
  • Keep humans on credentials and approvals. The AI did the work. Access stayed with me.

Want the longer version of this story?

I can walk you through the decisions, the trade-offs and what I would do next.