
{ESSAY} ·
07 MIN READ
↩ BACK TO JOURNAL
WHY MOST
REDESIGNS FAIL.
{POST/01}

{FIG. 01}
MATTIS REBRAND · 2024 → 2026
BEFORE / AFTER · CASE STUDY
Most teams treat the redesign as the goal. They ship the new homepage, post it on Twitter, and quietly let the numbers slide back to baseline within ninety days. The work was real. The change wasn't.

BY MARA VOSS
CREATIVE DIRECTOR
SHARE:
X
in
URL
Prompts define behavior, not configuration.
One evening last November, our AI assistant flagged a routine customer query as a potential exploit — and responded accordingly. Minutes later, alarms blared. Shortly after, we reverted the problematic prompt. Soon after, we found ourselves in a heated debate on how to prevent future incidents.
We had a prompt repository, vigilant reviewers, and detailed logs. Yet, we lacked a reliable method to foresee if a prompt would misfire. We were managing prompts like mere settings: easily tweaked, tracked in Git, peer-approved, and rapidly deployed. That Friday's disruption changed our perspective.
Prompts define behavior, not configuration.
Tweaking a YAML setting yields predictable outcomes. Prompts are nuanced. A minor adjustment can alter the model's attitude, its inclination to decline requests, its hedging behavior — across countless scenarios, subtly, beyond human grasp. Treating this as a config file was a fundamental error that became clear only after it caused problems.

“We were assessing prompts as if they were CSS tweaks. The model interpreted them as a complete overhaul of our user agreement.”
What we built instead.
We transitioned from a registry to version-controlled evaluation datasets — compact collections of inputs and anticipated outputs, each linked to a specific prompt version. Instead of code-review approvals, changes are validated by assessing the dataset and confirming that the new behavior aligns with the intended outcome. The comparison shifts from text to behavior.
The dataset resides alongside the prompt in the repository. CI executes it with each modification. Should a contributor introduce a prompt that unexpectedly alters rejection rates, the build fails — well before human review, and certainly before any user is affected.
Key takeaways.
Datasets must be agile.
If your evaluation suite requires hours to complete, it will be ignored locally. Aim for a focused set of examples per behavior, and actively remove examples that no longer expose genuine issues.
Production unveils the toughest cases.
While synthetic datasets offer comfort, they often miss critical failures. Regularly extract anonymized, deduplicated, and labeled slices from production logs.
Binary results are insufficient.
Monitor 'drift' as a key metric. A prompt that suddenly diverges from its previous behavior is noteworthy, irrespective of individual pass/fail outcomes.
Months later, we've eliminated late-night crises. We've reduced prompt changes — partly due to higher standards, partly because many prior changes were inconsequential. Validated changes deploy faster, with greater assurance, and remain stable.
Production AI requires rigor, not reinvention.
#evals
#prompt-engineering
#production-ai
#lessons-learned
GET WEEKLY INSIGHTS
One email every Tuesday.
AI strategies and studio updates. Join 12,000+ industry leaders.
