# Mohsin Amjed > Design and product executive working at the intersection of design, product, business, and AI Canonical site: https://mamjed.com Last updated: 2026-08-05 # About Source: https://mamjed.com/about ## Identity I sit between design, product, and business decisions, and I am most useful when the brief is half-written and the path forward still has to be argued for. Teams call me in when the product has outgrown its shape, the function needs structure, or the next bet has more risk than answers. I do not believe craft and strategy are different jobs. They are the same job, sequenced. ## Career arc Twelve plus years, six companies, one pattern. **Microsoft (2012 to 2013).** App Experience team. Cut global partner onboarding by over 60%. Learned how design lives inside a platform. **Samsung Electronics (2013 to 2015).** Co-founded Samsung Ads as the founding design lead. The business unit cleared $20M in profit in its first year. **SalesforceIQ (2015 to 2017).** Principal designer on Contact Gallery. The redesign lifted daily active users by 40% and duplicates merged by 34%. **Salesforce (2017 to 2020).** Senior product designer on Salesforce Essentials. Shaped SMB packaging and onboarding inside the bigger org. **Convoy (2020).** Principal designer on the SMB booking journey. Cut steps by 35% and lifted conversions by 18%. Piloted dual-track agile that raised design-to-development throughput by a quarter. **Sitetracker (2021 to 2023).** Director of Product Design and Head of Design. Built the function from zero: hiring, ladders, critique culture, the operating model the team kept running after I left. **Axios HQ (2023 to 2025).** Sr. Director of Product and Design. Re-architected the platform in three months and helped double AI usage across the product. The shape of the work is consistent. The brief is unfinished. The structure under the work has to be built before any screen earns its place. I move between vision and execution without losing either, because the best people I have worked for did the same. ## What I'm building now Three ventures in parallel, all making the same point: a senior design leader can hold craft, design operations, and AI fluency in the same pair of hands without losing any of the three. - **Nibbble.io.** Multi-tenant restaurant loyalty SaaS. Live in production at app.nibbble.io since April 2026. Three role-specific portals (admin, staff, customer), Square POS ingestion, Stripe billing, and a customer council of four restaurants pressure-testing the work. When most of the code is written by AI, the component library is the design system decision, so I chose HeroUI v3 for its agent-native artifacts. - **Simple Cortex.** AI consulting practice for Northern Virginia businesses and government. Landing live at simplecortex.com. The bet is that an operating company can run as a team of agents you can govern: persona, instructions, execution, and tool surface kept independently reviewable. Three agents committed to repo, each defined by a four-file contract. - **Kintsu Medspa.** Brand, voice, and website for a physician-led boutique medspa serving skin of color. I authored the brand with the client in person, then encoded that bar once so the system enforces it everywhere downstream. The same source produces the design in the codebase, the marketing playbook, the SEO audit, and the site plan. I authored 209 of 214 commits on the site. The sequence is the point. The brand and the design get built by hand. The systems that scale and codify the work get built with AI. The handoff is in the repo, not in a deck. ## How I lead The leadership style is practical, not stylized. I create clarity, protect the team, and raise the bar through critique, coaching, and decisions written down. I hire ahead of need, build the function around the people, use Jobs to Be Done as the shared decision language, apply design ops to AI contribution, and stay player-coach by default so I still ship. Culture is not separate from execution. Culture is how execution happens. [How I build a design function →](/leadership) ## AI fluency The proof is the current work, not the next essay. I treat AI as an operating model, not a novelty: the goal is to let work move faster while product decisions, brand rules, documentation, and review loops stay connected. Brand voice is enforced, not hoped for. Cost is made visible so the team can govern it. Customer councils pressure-test decisions before they ship, and bets carry kill criteria with public postmortems when they fail the gate. In practice that runs across three production repos, each carrying its own AI contract. I am less interested in AI as theater and more interested in AI as product infrastructure. The visible features matter. The work that quietly removes friction matters more. [How I work with AI](/ai-fluency) ## Personal context I am a faith-rooted operator. The same patience that lives in a critique room shows up in the community work. I serve as VP of Design and Innovation for the Ahmadiyya Muslim Youth Association (AMYA), the youth arm of the Ahmadiyya Muslim Community in the United States. Across the years, our chapters have raised and donated over $100,000 toward hunger relief and helped feed over 700,000 hungry Americans through coordinated food drives and community kitchens. That work shapes how I lead the day work too. Long horizons, small teams, real people downstream. ## Currently open to Director of Product Design, Sr. Director, VP, and Head of Design roles at small to mid product-led companies where a player-coach still matters. Advising and fractional work too, when the brief is specific and the stakes are real. If that sounds like the shape of your problem, [get in touch](/contact). # Leadership Source: https://mamjed.com/leadership ## What I believe My job is to build the conditions where design decisions get made well: clear direction from the top, real protection from organizational noise, a critique ritual that holds the bar honestly, and artifacts the team can run on after I leave. I lead by building the operating system, not just the team. The ladder, the hiring loop, the critique cadence, the JTBD frame, the readiness bar. A hiring manager inherits those when they hire a leader at this level. They are the work. ## Track record The pattern across five companies and three current ventures is the same. The function is stronger when I leave than when I joined, and the artifacts that made it stronger stay behind. Since Axios HQ, AI has written code, drafted copy, run audits, and caught regressions on every project I ship. - **Three ventures, all currently shipping, AI in the production loop on each.** Nibbble.io in production with an unattended local-model autofix loop. Kintsu Medspa with a governed AI workflow that designs the product, writes the marketing, and audits the growth, against a bar committed once. Simple Cortex running as a three-agent company with cost as a tool surface. (2025 to present) - **Re-architected the Axios HQ product in 90 days. Doubled AI usage across the platform.** The real problem was not navigation. The product had outgrown its original shape and every new feature was fighting for room. I led the multidisciplinary product organization through the rebuild and shaped AI strategy with the CMO and Head of AI. (Axios HQ, 2023 to 2025) - **Built the Sitetracker design discipline from scratch.** Skill matrix, spec-free hiring loop, career ladder, mentorship model, critique culture, and a global capability in India that ran without me in the room. The team had good people. It did not yet have the conditions for design to shape product decisions. (Sitetracker, 2021 to 2023) - **Co-founded Samsung Ads. $20M in profit in year one.** A small team with a large company's surface area. I was the first UX designer on the unit and led design as it grew into a business. (Samsung Ads, 2013 to 2015) - **SalesforceIQ Contact Gallery: +40% DAU, +34% duplicates merged.** Shipped against a live CRM with measurable lift. The early AI UX work on confidence, source transparency, and correction loops fed forward into Salesforce Einstein. (SalesforceIQ, 2015 to 2017) - **Cut Microsoft global partner onboarding time by over 60%.** Early proof that systems thinking pays at platform scale. (Microsoft, 2012 to 2013) ## How I build a design function Six artifacts. I install them when I take a function from zero, or when I rebuild one that has drifted. The first five have held their shape across three companies. The sixth is new in the last two years and is already load-bearing. **Skill matrix.** Capabilities mapped across a small set of named dimensions. Gaps tracked openly. Hiring, promotion, and coaching all reference the same matrix, so growth conversations are not reinvented per person. AI fluency lives on the matrix as its own dimension. Not a bonus. **Spec-free, skill-based hiring loop.** I do not hire from a job description. I hire against the matrix and the gap. Roles are scoped from the team's actual capability shape, not a generic ladder cell. The interview asks what the person will do in their first ninety days, not which credentials they hold. **Career ladder.** Levels and expectations written down, visible to the team, referenced in every career conversation. The ladder is a contract. It tells a designer what next looks like and what evidence proves it. **Critique culture.** Critique is a standing practice, not a vibe. Designers present problem context before they show the design. Notes tie back to a user need or a business goal, not aesthetics. A safe team is not a comfortable team. It is a team willing to challenge the work. The bar is consistent across shipped work, prototypes, reviews in progress, and anything an AI helped produce. **Jobs to Be Done as the organizing primitive.** Roadmaps, research, packaging, and design briefs anchor to the job, not the feature list. At Sitetracker, JTBD became the shared vocabulary across product, engineering, and design. It outlasted my tenure. ## Operating rituals **One-on-ones** keep status out of the conversation. Status lives in the doc and the standup. This time belongs to the designer's career arc, the work that is hard, and what they want pushed back on. The ladder enters every sixth week so growth is never abstract. **Critique** is the readiness mechanism. Work that survives critique is work ready to ship. Written prep is expected before anyone opens a file: the problem, the constraints, the open questions. Notes tie to a user need or a business goal, not aesthetics. AI-assisted output is held to the same bar. **Design and product sync** puts design at problem framing, not handoff. At director scope, the question is whether the work answers the right Job to Be Done and whether the team has earned the next decision without me in the room. **Quarterly portfolio review** is the receipt the leadership team gets when budget conversations come around. Three slides per designer. The artifact lives in a shared doc so the whole team can read each other's arc. **Quality bar, written down.** Readiness is not a feeling. It is a checklist the team co-authored and enforces on each other. Brand voice, accessibility, performance budgets, evidence of testing, and the AI contribution contract are all line items. ## When the framework leaves the design org At Sitetracker, the first move was not a workshop. It was a set of conversations with the Head of Product and engineering leadership about what was making product decisions slow. The answer, consistently, was that product, design, and engineering were starting from different assumptions about the customer. Not different data. Different mental models of what the customer was trying to accomplish. Jobs to Be Done entered the room as a shared frame, not a design deliverable. Over several months, it became the organizing logic for how product conversations started: with the job, not the feature. The Head of Product embedded it in the PDLC playbook. Engineering leads started asking "whose job is this solving?" in sprint reviews. The design team did not own JTBD. The org did. That is the part that mattered. The framework outlasted my tenure because it was never positioned as a design methodology in the first place. ## The twelve-month bar I measure leadership by what people become, not just what they ship. The coaching is concrete: clearer writing, sharper presentation under pressure, the ability to run their own critique, the confidence to own a career conversation, and the judgment to use AI as a thinking partner without dropping the bar on their own craft. The artifacts make the growth legible. The coaching makes it real. The pattern I look for is a designer who, twelve months in, no longer needs me to translate the business context for them. That is the proof. The designers I have managed and the partners I have worked with across product and engineering describe the practice more directly than I can. > "What sets Mohsin apart as a mentor and leader is his ability to teach a team the necessary skills to solve complex problems instead of being prescriptive with a solution. This not only contributes to the delivery of a high quality product, but it grows the skillset of each member of the team." > > Bailee Warsing · Sr. Product Designer, Sitetracker · Direct report > "Mohsin brings the much-desired balance of both theory and practice, stemming from many years of hands-on experience. Couple that with his strong empathy, deep caring, and absolute passion for design and it is no wonder Mohsin is able to attract, build, and grow top-notch design teams." > > Ian Ezra · Head of Product, Sitetracker · Cross-functional peer > "Of his many talents, the thing I appreciated most about working with Mohsin was his ability to give constructive criticism in a motivating way. With Mohsin, everything boils down to better serving the customer, which reduces politics and infuses focus and purpose to every project." > > Alexis Smith McFarlane · Sr. Product Manager, SalesforceIQ · Cross-functional peer [How I work with AI in detail](/ai-fluency) # AI Fluency Source: https://mamjed.com/ai-fluency # AI Fluency > AI fluency · Current practice **AI as product infrastructure, not theater.** Visible where it earns the user's trust. Invisible where it earns their time. That call sits with me, not the model. At Axios HQ I led the AI strategy that doubled platform usage. Earlier at Salesforce I shipped some of the first Einstein experiences. Today the practice runs in production across three ventures: Nibbble, Simple Cortex, and Kintsu Medspa. Most AI features do not earn their place. The ones below do. ## Shipped, not theoretical - **Axios HQ** · **2x** · AI usage across the platform. Sr. Director of Product and Design, 2023 to 2025. - **Nibbble autofix** · **17** · Real production fixes shipped by an AI assistant working unattended on a small, fenced-off list of files. The local model runs free per attempt. - **Nibbble repo** · **412** · Code commits to the live product in twelve weeks. Mohsin wrote most of them. Claude paired on nearly half. - **Nibbble coverage** · **95.75%** · Automated tests cover almost every line of a live SaaS product serving multiple restaurants. 21 architecture decisions are written down in the repo. - **Kintsu brand system** · **2,745** · Lines of brand strategy I wrote and made AI-enforceable, across six long-form documents AI agents read on every session. ## Where it shows up in the work ### Axios HQ Re-architected the product in 90 days and doubled AI usage across the platform. Sr. Director of Product and Design. [Read the case study →](/case-studies/axios-hq) ### Nibbble.io A loyalty product serving multiple restaurants, live since April 2026. Code my AI assistant wrote, reviewed, and committed while I slept: 17 real fixes in production, each scoped to a small list of files I pre-approved. The codebase is mapped into a 3,339-node knowledge graph the assistants check before they search, so they understand the system before they touch it. [Read the case study →](/case-studies/nibbble) ### Simple Cortex Three AI agents that run a consulting practice. Each agent's job description, instructions, runtime behavior, and tools live in their own file in a repo, so anyone can review what an agent is supposed to do separately from how it does it. [Read the case study →](/case-studies/simple-cortex) ### Kintsu Medspa [Read the case study →](/case-studies/kintsu-medspa) ### Deep-researcher skill A custom research assistant I built. Takes a brief, goes off and researches, returns sourced findings. Modeled on Manus.im and built on top of Hermes, the open-source agent from Nous Research. Plan, Act, Observe, Update, with citation tracking and local-first model routing. It produced the Kintsu marketing playbook, the SEO and AEO audit, and an HRT business-intelligence report under brief. Operator note ## Working principles > Not aspirations. The rules my repos and my teams run on today. How I work with AI ### 01. Visible where it earns trust. Invisible where it earns time. AI shows up in the interface when seeing it helps the user. It stays out of the way when it is just a faster path to the same answer. That call is the same on every product surface, and it is mine, not a model's. ### 02. Local-first routing. Cloud where judgment matters. Most AI tasks are cheap and bounded. Those run on free, local models. The hard, irreversible calls (architecture, judgment about user trust) go to Claude in the cloud. Cost shows up where the decisions get made. ### 03. Voice and brand as contracts the AI reads every session. Brand voice, prohibited words, motion rules, color tokens, and compliance rules live as files in the repo. Every AI agent reads them before it writes. Drift gets caught at the contract stage, not at review. ### 04. Cost and provider health are operating surfaces, not slides. I can check what every AI call cost me today the same way I check the weather. One command at the terminal. Provider health and per-route spend are everyday operator surfaces, not a quarterly review item. ### 05. Agents review agents. One AI proposes the work. Another AI reviews it before it ships. Codex blocks pull requests on the Nibbble repo if the change misses the bar. Humans merge. The AI accountability ladder is not optional. ## What stays out of the work ### 01. AI demo theater. No future-of-design essays. No speculative roadmaps. No logo grids of model providers standing in for capability. ### 02. AI features without a fallback. Every AI surface ships with an explicit handoff to a human and an editable output. ### 03. Speed paid for by dropping the bar. Cycle time and quality get measured side by side. AI shortens the path. It does not move the destination. ### 04. Final calls delegated to models. Product direction, user trust, hiring, team direction. Those stay with humans. ## The stack **One brain. Many models. Cost on every call.** Agents ask for a capability, not a model by name. A router picks the right model, falls back if it has to, and tracks what the call cost. Local models handle the cheap, bounded work. Cloud models handle the judgment calls. ### Layer 01 · Surfaces Where the work happens. Each surface reads the same repo-level contracts. - **Claude Code** · Primary engineering pair - **Codex CLI** · PR review gate - **Gemini CLI** · Research and polish - **Warp** · Agentic terminal - **Conductor** · Parallel CC sessions in tabs - **Perplexity** · Live research surface - **Manus.im** · Autonomous research and writing service - **Claude Agent SDK** · Programmatic agents ### Layer 02 · Routers Agents call by intent. Routers classify, fall back, and meter cost. - **OpenClaw Router** · Classify · architect · engineer · analyst · triage - **Hermes** · Semantic router across skills and providers - **Custom routers** · Per-project intent-to-model mapping ### Layer 03 · Providers Cloud where judgment matters. Local and hosted open source for bounded structure. - **Anthropic** (Cloud) · Cloud judgment - **OpenAI** (Cloud) · Codex review - **Google** (Cloud) · Research - **Alibaba Cloud** (Cloud) · Coder plan, Kimi access - **Ollama** (Local) · Local runtime - **Fireworks** (Hosted OSS) · Hosted OSS inference ### Layer 04 · Models Named, not aspirational. The exact models the agents resolve to today. - **claude-opus-4-6** (Cloud) - **claude-sonnet-4-6** (Cloud) - **gpt-5** (Cloud) - **gpt-5-codex** (Cloud) - **gemini-3-pro** (Cloud) - **gemini-3-flash** (Cloud) - **kimi-k2-5** (Cloud) - **kimi-k2-6** (Cloud) - **qwen3.6:27b** (Local) - **qwen3-coder-next:q4_k_m** (Local) - **openclaw-qwen35-a3b-think** (Local) - **qwen3.5:9b** (Local) - **qwen3.5:4b** (Local) - **llama3.2:3b** (Local) ### What runs unattended An overnight job on the free local tier tunes routing and prompts on its own. Thirty dated research branches over two months. The cheap tier earns its keep. ## Hiring a design leader who actually ships AI? - [See the work](/case-studies) - [Hire me →](/contact) # Resume Source: https://mamjed.com/resume # Resume **Mohsin Amjed** Design and product leader. I build the teams, the platforms, and the clarity that lets a product earn its next phase. Northern Virginia, USA - Email: [mohsin@mamjed.com](mailto:mohsin@mamjed.com) - LinkedIn: [linkedin.com/in/mamjed](https://www.linkedin.com/in/mamjed/) - Site: [mamjed.com](https://mamjed.com) --- ## Summary Twelve-plus years across Microsoft, Samsung, Salesforce, SalesforceIQ, Sitetracker, and Axios HQ. Built design functions from zero. Re-architected platforms under pressure. Shipped measurable outcomes against AI-enabled surfaces. Currently building three ventures in parallel: Nibbble.io, a multi-tenant restaurant loyalty SaaS in production; Simple Cortex, an AI consulting practice running on a three-agent operating company committed to git; and Kintsu Medspa, the brand system, voice, and website for a physician-led practice. Operating mode is player-coach. I author the brand, set the voice, codify the AI contribution bar in the repo, and hold the merge button until the team is ready to take it. --- ## Currently building - **Nibbble.io.** Founder and builder. Multi-tenant restaurant loyalty SaaS, in production since 2026-04-22. Square POS ingestion, Stripe billing, and an AI-assisted engineering loop. See [/case-studies/nibbble](/case-studies/nibbble). - **Simple Cortex.** Founder and principal. AI consulting practice for Northern Virginia businesses and government. Runs on a three-agent operating company committed to git. See [/case-studies/simple-cortex](/case-studies/simple-cortex). - **Kintsu Medspa.** Brand, voice, AI tooling, and technical owner of the website build for a physician-led practice serving skin of color. Grand opening end of May 2026. See [/case-studies/kintsu-medspa](/case-studies/kintsu-medspa). --- ## Experience ### [Axios HQ](/case-studies/axios-hq) **Sr. Director of Product and Design** March 2023 to June 2025 - Doubled AI usage across the platform. - Re-architected the product information architecture in 90 days. - Led the multidisciplinary product and design org through a period of AI-driven product change. - Set the planning cadence, research practice, and operating model the team ran on. - Shaped AI strategy with the CMO and Head of AI at executive altitude. ### [Sitetracker](/case-studies/sitetracker) **Director of Product Design, Head of Design** June 2021 to March 2023 - Built the design discipline from scratch: hiring loop, career ladder, critique culture, mentorship model. - Made Jobs to Be Done the shared decision-making frame across product, engineering, and design. - Connected design to product and engineering on a weekly cadence the team kept after I left. - Hired and coached the founding design team across staff, senior, and mid-level designers. - Stood up a global capability in India that scaled the team without dropping the bar. ### Convoy **Principal Designer** June 2020 to November 2020 - Streamlined the SMB booking flow: 35% fewer steps, 18% lift in conversions. - Piloted dual-track agile. Increased design-to-development throughput by 25%. - Prototyped a dynamic pricing UI that lifted average booking value on day one. ### [Salesforce](/case-studies/salesforce-essentials) **Senior Product Designer, SMB and Salesforce Essentials** April 2017 to June 2020 - Shaped Salesforce Essentials packaging and onboarding for the SMB segment. - Translated SMB user behavior into product framing, packaging, and adoption decisions. - Bridged design, product, engineering, and GTM through launch and iteration. ### [SalesforceIQ](/case-studies/salesforce-iq) **Principal Product Designer** December 2015 to April 2017 - Led the Contact Gallery redesign: +40% DAU, +34% duplicates merged. - Owned product design on the relationship intelligence surface. - Shaped early AI UX patterns (confidence signals, source transparency, correction loops) that fed forward into Salesforce Einstein. ### [Samsung Electronics](/case-studies/samsung-ads) **Head of Design, Samsung Ads** November 2013 to December 2015 - Co-founded Samsung Ads inside Samsung Electronics. Led design from zero. - $20M in profit in the unit's first year. - Shaped the product surface and the ad-format work that opened the business. - Contributed to early IAB standards thinking for connected TV advertising. ### [Microsoft](/case-studies/microsoft) **Product Designer, App Experience Team** July 2012 to November 2013 - Cut Microsoft global partner onboarding time by over 60%. - Designed app experiences across the partner ecosystem. --- ## Community and volunteer leadership ### MKA USA (Ahmadiyya Muslim Youth Association) **VP of Design and Innovation** January 2013 to present - AMYA raised and donated over $100,000 toward hunger relief. - AMYA helped feed over 700,000 hungry Americans. - Lead design and innovation across national volunteer programs. --- ## Skills and competencies - Design leadership and team building - Product strategy and platform thinking - Information architecture under pressure - Zero-to-one product design - Design systems and design ops - AI-enabled product surfaces - Jobs to Be Done research practice - Cross-functional operating models - Brand systems and voice as a system - Player-coach craft: visual design, prototyping, and shipping # Contact Source: https://mamjed.com/contact # Get in touch The clearest way to reach me is email. I read everything that comes in, and I prefer messages that are specific over messages that are polished. If you are hiring, building, or thinking out loud about a design org or an AI product, I want to hear about it. ## What I am open to - Senior design and product leadership roles. - Advising and fractional engagements with seed, Series A, and Series B teams. - AI product strategy consulting via [Simple Cortex](https://simplecortex.com). - Speaking, panels, and podcast invitations on design leadership, AI-native product work, and operating models. ## How to reach me - **Email.** [mohsin@mamjed.com](mailto:mohsin@mamjed.com) - **LinkedIn.** [linkedin.com/in/mamjed](https://www.linkedin.com/in/mamjed/) - **Booking.** [Book a 30-minute call on Calendly](https://calendly.com/mamjed/30min). Based in Northern Virginia, USA. I work on US Eastern time. ## What to include in the first message A short note with these four lines saves us a round trip and tells me whether I am the right fit: - **Company and stage.** Name, stage (seed, Series A, Series B, public), and a one-line description. - **Role or engagement.** Full-time, fractional, advisory, consulting, speaking, or "still thinking." - **Timeline.** When you want to talk, when you want to decide, and when the work would start. - **What brought you here.** A case study, a recommendation, a referral, or a specific question. This helps me reply with something useful instead of generic. ## Response expectation I respond within one to two business days. If you do not hear back within a week, resend. Email gets buried, not ignored. ## What I am not the right fit for I am not the right fit for pure visual or marketing-design contracts, agency-style production work, or roles that want a designer to execute against a fixed spec without a seat at the strategy table. If that is the brief, there are people I can refer you to. # Case studies ## Axios HQ Source: https://mamjed.com/case-studies/axios-hq # Axios HQ ## The short version Sr. Director of Product and Design at Axios HQ, leading an eight-person org across information architecture, product design, research, analytics, and AI product strategy. The proof: AI usage doubled across the platform. I rebuilt the foundation in the first 90 days, matured product, design, research, and analytics into one organization, and set the AI strategy with the CMO and Head of AI while engineering owned delivery. ## The product had lost a shape anyone could hold Navigation was the symptom. Axios HQ started as a focused writing tool and became a platform in waiting. Every new feature fought the old structure before it could create value. The constraint was the mental model itself. The product no longer had a shape anyone could hold in their head, and the company was about to ask it to carry a much bigger roadmap. I rebuilt the information architecture first because planning, research, collaboration, measurement, and AI all depended on it. ## What I did - Re-architected the product experience in the first 90 days so the AI work and feature roadmap that followed had a coherent structure to live inside, not bolted onto an app that no longer made sense. - Chose a series-and-channels model over a folder hierarchy for the core structure. Folders were the obvious path and the one the flat newsletter list already implied, but folders model storage, not editorial intent, so every new feature had to renegotiate where it lived. Series and channels modeled the actual work and the roadmap. I made the call to restructure around the work, and the roadmap stopped fighting the structure. - Led product, design, research, and analytics as one organization: built and matured each function in place rather than running them as separate workstreams. - Built research infrastructure (personas, JTBD frameworks, feedback loops) as durable assets the team could draw on continuously, not one-off studies archived after a readout. - Introduced Shape-Up-inspired planning and North Star vision framing so the team could commit to bets without losing executive visibility. - Set the AI product strategy with the CMO and Head of AI, landing "visible versus invisible AI" as a decision rule: AI that earns user trust surfaces itself, AI that earns user time stays out of the way. Making it a rule rather than a review meant teams could resolve most AI calls themselves instead of escalating each one. The payoff was that AI stopped being a press release and became a product principle. The visible-versus-invisible frame guided individual product decisions without executive arbitration on each one, and platform-wide AI usage doubled. That was the structural bet paying off: the foundation was what let the AI work land as a native layer instead of a bolted-on feature. What changed, beyond the AI number: - Research went from reactive studies to standing infrastructure. Personas, JTBD frameworks, and feedback loops became durable assets the team drew on continuously. - Planning became something the team owned. Shape-Up-inspired bets and North Star framing let the team commit without losing executive visibility. ## Reflection 1. Pull research before the first architecture pass, not after. I committed the first IA cut partly on gut, then reworked decisions once the mental-model research caught up. Starting from user mental models would have sharpened that first pass and saved the downstream iteration. That is the honest cost of moving fast on structure. 2. A decision rule scaled better than a decision review. Turning "visible versus invisible AI" into a rule teams could apply themselves, instead of a call I arbitrated case by case, is what let the AI work move without me in the middle of every judgment. It is the lever I would reach for first on the next platform. ## Kintsu AI Consultation Tool Source: https://mamjed.com/case-studies/kintsu-ai-consultation-tool # Kintsu AI Consultation Tool **Shipped to production July 2026. Live at imagine.kintsuaesthetics.com (noindex pending attorney and Medical Director sign-off on BAA determination).** ## TL;DR A selfie-based AI consultation only earns its place if it earns trust first. Running the Simple Cortex engagement for Kintsu, I designed an ephemeral flow that reads a photo, maps concerns to the real treatment catalog, and moves a patient toward booking. No stored photo, no database record, no login. It is ephemeral by architecture: the photo never touches disk, so there is nothing to leak. Shipped to production in July 2026. The proof is the architecture first, and the live system second. ## The problem was not AI accuracy Accuracy matters, but trust comes first. The tool asks for a sensitive image: a person's face. The first design decision was not which model to use. It was whether this product deserved to keep the photo at all. The answer was no. So the tool became ephemeral by architecture, not by policy. Take the image, analyze it, return useful guidance, map results to services the practice actually offers, then get out of the way. A privacy policy saying "we delete photos later" was not good enough. The product promise is stronger when there is no stored photo to delete. Useful enough to convert, restrained enough to trust. The second constraint was clinical: the line between patient education and medical advice can blur quickly in a medspa context. A tool that casually overclaims, or one that treats a light skin tone as the default and applies the same confidence to darker skin, damages trust before the consultation starts. I designed around both risks from the first conversation. ## What I did - Designed a selfie-based consultation flow with no persistent photo storage. The image travels in the request, passes through the model, and the result returns to the browser. No write to disk, no row in a table. - Mapped concerns to the actual treatment catalog rather than generic advice. The model can describe a concern. The recommendation has to come from what Kintsu actually offers, so the service catalog acts as a guardrail (see the decision beat below). - Created a plain-language result screen that educates without pretending to diagnose. The tool says what it notices and what people with similar concerns often explore. It does not say what condition a patient has. - Built guardrails into the system prompt itself: no diagnostic language, skin-tone-aware behavior required, recommendations limited to the vetted catalog. The system prompt is the design artifact. - Framed the tool as a conversion bridge to booking, not a replacement for clinical judgment. The booking handoff carries the patient's shortlist of concerns and services, never patient identifiers. The user gets a useful next step. The business gets a lead path. The clinician keeps authority. - Isolated the AI image simulation behind a feature flag. The useful consultation flow ships without waiting on the legally exposed feature to clear review. - Ported the API from Vercel serverless to a Hono Node server compiled with esbuild, deployed as a Docker container behind Coolify's Traefik on the Hostinger VPS. Auto-TLS, rightmost-XFF trust, and a noindex header ship as infra configuration, not application code. - Replaced the single-shot selfie with a 5-frame white-light capture protocol: ambient, full-flash, left-raking, right-raking, and top-band. Raking shadows are real geometry physics a screen can produce; that reasoning ruled out colored flashes (screens cannot fake UV or polarized physics). Consent v2 adds a photosensitivity warning and a single-photo opt-out path. A quality gate retries up to three times and surfaces the failure reason as visible text rather than a silent stop. - Hardened the server boundary: Cloudflare Turnstile wired fail-closed (the secret lives only on a managed siteverify Worker so the token is never pre-verified on the frontend), rate-limiter sweep and cap, intake enum validation to close prompt-injection, CSP, and a 413 pre-buffer. Dependencies brought to zero known vulnerabilities. - Wired Zenoti lead integration: email-then-phone deduplication before guest create (fail-open, a duplicate row beats a lost lead), guest note attached with the correct field contract (field name, note type enum, and center as a nested object, not a flat ID). A code-review pass caught HTML injection in the front-desk notification email and added escaping plus an injection test before deploy. **The decision beat: I made storage impossible, not just discouraged.** The reflexive engineering path was to write each analysis to a row so we could debug it later, backed by a privacy policy promising to delete photos on a schedule. I ruled that out before it was built. A policy can be reversed, breached, or quietly ignored. The stronger promise is architectural: if the photo is never written to disk or a table, there is no stored photo to leak, subpoena, or forget to delete. So the flow was ephemeral by architecture from the start. The cost is real: I gave up server-side debugging of individual analyses, which makes some failures harder to reproduce. I took that trade because on a tool handling faces, "we cannot leak what we never kept" is worth more than convenient logs. ## What changed The tool shipped to production in July 2026. What changed is both the shape of the product and what the real deployment surfaced. - The design moved from a risky AI novelty to a bounded experience: capture, analyze, return, clear. There is no persistence layer for the photo, so the riskiest failure mode is designed out rather than mitigated after the fact. - The recommendation surface is closed, not open. Instead of trusting the model to name a treatment, the catalog guard checks every recommended service against the live catalog and drops anything the practice does not offer, before the response leaves the server. That is a check I can point to, not a claim to take on faith. - Darker skin tones are treated as a first-class case in the system prompt, not covered by a disclaimer after the fact. - The booking handoff carries the patient's shortlist of concerns and services, never patient identifiers. - The AI image simulation stays behind a feature flag until legal and clinical review clear it, so the useful flow ships without waiting on the exposed feature. - The production launch surfaced a latent model configuration bug: a prompt-parameter fix activated a previously-inert `maxOutputTokens` cap on Gemini 2.5 Pro, producing "No object generated" errors on live requests. The fix that improved the prompt broke the model in a way local tests could not catch, because the cap only triggers at inference scale. Production was pinned to GPT-4.1 while the token budget was diagnosed and raised. The enhanced error logging now captures `finishReason` and `usage` from the AI SDK error object so any future truncation is self-diagnosing. **Hypothesis (pre-patient-facing, pending BAA clearance):** a trustworthy middle step between curiosity and booking should raise the rate at which visitors schedule a consultation, because they arrive with a concrete shortlist and a reason. I have not measured this. Treat it as the thing to test at launch, not a result. ## Reflection 1. **Restraint was the product, not a constraint on it.** The tool earns trust by refusing to keep the photo, refusing to diagnose, and refusing to recommend anything outside the real service model. Each refusal made the thing safer to launch and easier to trust. On a product handling faces, the strongest feature is often the data you choose not to hold. 2. **Ephemeral by architecture cost me observability, and I would make the same call again.** Because no analysis is written to a row, I cannot replay an individual bad result to see what the model saw. That is a real debugging tax. I accepted it because the alternative, a store I promise to delete, trades a hard guarantee for a soft one. The open question I am carrying into launch: can I add aggregate, non-identifying quality signals without reintroducing a photo store through the back door. 3. **A fix that improves the prompt can break the model.** The Gemini outage was caused by a parameter correction that activated a token cap only reachable at real inference scale. The lesson is not to distrust AI SDK configuration, but to treat model behavior in production as a separate test surface from local runs: instrument `finishReason` and `usage` from the start, because truncation at scale looks identical to a correct result until someone notices the output is empty. ## Kintsu Medspa Source: https://mamjed.com/case-studies/kintsu-medspa # Kintsu Medspa ## TL;DR Running a Simple Cortex engagement, I built a physician-led medical aesthetics brand from a blank slate and turned it into an operating system for voice, imagery, marketing, SEO/AEO, and AI-assisted contribution. I held the pen alone: I wrote the six-document brand system and authored 209 of the 214 commits in the site repo. Live at kintsuaesthetics.com since January 2026. The proof so far is internal but real: the brand compressed into a 71-line contract AI contributors read every session, and the site shipped with an honest 79/153 (52%) SEO/AEO baseline and a plan to reach 85%+ by Q1 2027, not a vanity number. ## Polish is cheap. Trust is the hard part. A new medical aesthetics brand can look polished and still say nothing. The harder job was building trust before the first appointment. The category is full of soft-focus luxury language, generic transformation promises, and sites that treat skin of color as an afterthought. This practice was physician-led, clinically careful, and built for patients who have often been underserved by aesthetic medicine. I treated the brand as infrastructure, not a moodboard. Clinical without cold, warm without vague, premium without sounding like every other medspa. And the system had to hold once AI-assisted contributors joined the production workflow, not just while I held the pen. ## What I did - Built the brand from a blank slate: voice, imagery direction, marketing architecture, AI design guidelines, SEO/AEO audit, and site plan, delivered as six long-form documents in the repo. - Set five voice pillars with the client face to face before the site system hardened. The tagline "Repair. Restore. Radiate." only works if the voice earns it. - Designed for skin of color as the default patient, not an addendum. Positioning, treatment education, and page hierarchy make Fitzpatrick III through VI care visible before the first scroll. - Encoded the bar where the work happens: brand guidelines, voice rules, prohibited words, and AEO guidance written as living artifacts the build could use, not PDFs outside the product. - Ran a 153-item SEO/AEO audit and set the baseline honestly at 79/153 (52%), then wrote a maintenance plan targeting 85%+ by Q1 2027 instead of shipping a vanity number. The voice came down to one line I killed. The category defaults to soft-focus luxury copy, and the first draft of the treatment language reached straight for it: "Radiance, reimagined for you." It scans as premium and says nothing a patient can trust. So I cut it and wrote what should replace it: "Evidence-based care for skin that has been overlooked." One is aspirational and generic. The other is specific, clinical, and names the patient the category ignores. I could have kept the pretty sentence and moved on. Instead I turned the choice into a rule, a SAY THIS / NOT THIS pair in the Voice and Imagery Guidelines, because a physician-led practice earns trust by being concrete, and because a rule survives handoff where a single good sentence does not. Then I mirrored the same pairs into the AI contributor contract, so the next time an AI drafts a page it inherits the judgment instead of reaching for the luxury cliché again. The line I killed is the reason the ones that follow it hold. The practice launched with skin-of-color positioning leading the story instead of sitting in a footnote, and the voice rules did not stay on paper. They run in the build, where any reviewer, human or AI, can cite a prohibited word at pull-request time. The point was never the document count. It was that a new contributor can execute the brand without re-litigating taste every time. That is design leadership applied to brand: make judgment portable. ## The booking portal The brand system was always meant to close the loop at the booking step. The second phase shipped a live `/book` funnel backed by a production API at `api.kintsuaesthetics.com`. - Built real treatment category tabs by mapping Zenoti's services API server-side (parallel fetches, 1-hour cache). Added provider-scoped deep links so a staff member can share `/book?p=pooja-shah` and the funnel locks to that provider's eligible services and live slots. Added a session-scoped client cache so returning to the time-selection step is instant instead of re-fetching availability. - Hardened the backend for production: retry with Retry-After honor, TTL caches at two tiers, body-key idempotency, per-IP rate limiting, strict CORS allowlist, and PHI-scrubbed logs. Deployed as a Docker container behind Coolify's Traefik reverse proxy, TLS-terminated at `api.kintsuaesthetics.com`. - Chose VPS over AWS after re-assessing the practice's HIPAA status: Kintsu does not bill insurance, so it is likely not a covered entity and a BAA is not legally required. That finding eliminated the operational overhead of AWS KMS, CloudTrail, and formal BAA paperwork while retaining the technical privacy controls already built. The decision and revert triggers are documented in ADR-0010; the AWS plan is preserved as a fallback. - Hardened the funnel for touch: `useHasHover()` gates hover interactions behind a pointer-capable device, so touch taps do not trigger or stick desktop hover states. Summary remove targets are always visible on touch (36px) and reveal-on-hover only on a pointer device. Measured tap targets in the browser before changing them; two speculative fixes were dropped because the defaults were already correct. ## Measuring it, and letting it run A brand system and a booking funnel are inputs. Neither tells you whether the money going out the door is working. Paid ads were running with no automated read on them at all, and the weekly digest designed to fix that had been written months earlier, never committed, and never once run. The third phase closed that gap: an experimentation program on the site, a monitor on the spend, and a deliberate answer to how much either is allowed to do on its own. - Shipped the A/B program in three phases (2026-07-25). Assignment goes through a permanent `getVariant()` facade backed by the GrowthBook SDK in local-payload mode, so call sites never change if the payload source does. Standing rule for the program is small maintained OSS over hand-rolled bucketing or statistical numerics. - Made the power discipline mechanical. Every experiment declares minimum sessions per variant (typically 1,500) and minimum duration (typically 28 days) at creation, and both must be met independently before anything can read as a decidable win or loss. The gate is a script with exit codes, not a judgment call, because "does this look significant yet" is exactly the question people answer wrong when they want a result. - Baked control into the prerendered HTML and made assignment return control whenever the browser reports itself as automated. One rule covers the Puppeteer prerender, the Playwright suite, and most crawlers, which means variants can never become accidental SEO cloaking. - Amended the concurrency rule in writing rather than silently. The program launched at one running experiment site-wide because the site runs around 38 sessions a day. ADR 0011 raised that to three, but only on distinct routes with independent traffic, specifically so a paid-campaign landing page would not wait behind an unrelated hypothesis. The per-experiment floors did not move, and the original decision text was preserved and annotated instead of rewritten. - Built the daily ads monitor with a deliberately asymmetric autonomy envelope (ADR 0012). The autopilot may pause an ad, pause an ad set, or decrease a budget. That is the whole allowlist. It can only spend less money and show fewer ads, every action is reversible in one click, and copy, creative, targeting, budget increases, and anything price-bearing stay permanently human. - Enforced least privilege in-process because the vendor offers no key scoping. Two independent conditions must both hold for any write, blocked writes throw before any network request, and every write emits an audit line into the immutable Actions log on both the success and failure paths. All scheduled work runs on an always-on Linux VPS runner, so nothing depends on a laptop being open. The decision I care most about here is what I kept out. No LLM sits in the decision path or the write path. The account produces roughly seven conversions a week with a one to three day lag, which is about one data point per day, and a loop that is always thinking against that stream is a loop reacting to noise it cannot distinguish from signal. The shape is three batch tiers instead: a daily deterministic sweep, a weekly digest, a monthly human review. Language models synthesize the narrative and generate hypotheses for a human. They never execute. I had a fresh-context architecture review run against the frozen design looking for the case that I was being too conservative, and it argued independently against an always-on agent. Stress-testing before enabling any of it surfaced four defects that all reported healthy. The guardrail protecting running experiments matched experiment ids against campaign names and scored zero matches across eight campaigns, so it was decorative; it now joins on the ad's destination URL. The zero-conversion alert used a share-of-spend rule that sat where the probability of a legitimate zero was 37 percent, which would have trained me to ignore it inside a week; it is now derived from the account's own CPA. One live campaign pointed at a hash-fragment booking URL that the pathname-routed site resolved to the home page. And two different budget situations, a capped winner and an efficient campaign that never reaches its cap, were being treated as the same finding. The honest outcome so far is that the first two hero experiments were stopped as unconcludable rather than called. At this traffic level the floors bite, which is what they are for. The system's value right now is that waste gets caught daily instead of whenever someone remembers to look, and that a winner cannot be declared early by anyone, including me. ## Reflection 1. The brand foundation is internal proof. The booking portal is now live and taking appointments. The next metric that matters is market behavior: search visibility climbing, consultation starts, booking conversion. The system is built to measure those. Until the numbers land, I should say so plainly rather than dress a baseline up as an outcome. 2. Voice drift is the predictable cost of putting AI in the production model. If I ran this engagement again I would pair the SEO/AEO audit with a quarterly voice review from the start. The answer is not less AI. It is a tighter review rhythm so the 71-line contract stays honest as the site grows. 3. I would bring a design systems engineer in earlier. I authored 209 of 214 commits, which made me the single point of failure on the build. That was fine for a launch and would not scale past one. ## Kintsu Portal Source: https://mamjed.com/case-studies/kintsu-portal # Kintsu Portal ## TL;DR I led this for Simple Cortex as product and design lead, from blank schema to opening day. I was the sole human author, with Claude Code as orchestrator. The proof I trust most. A blank schema became 15 deployed pages in 72 hours. 544 tests passed at the foundation merge. A PHI gate scans every migration and has caught zero PHI. The cockpit-first framing, the pricing-engine boundary, and the PHI-free architecture were my calls. The cockpit ships the same message the whole build carries: a daily surface for decisions, not another admin panel full of settings. ## A cockpit, not an admin panel Admin panels show everything. That was the wrong answer. A first-time practice owner does not need every object in the database exposed as a table. She needs to know what requires attention, what is ready, what could break before opening day, and what is unresolved that a decision could close. Running the Simple Cortex engagement, I designed the portal as a cockpit, not a control panel. The system has real complexity: pricing logic, service taxonomy, staff workflows, AI, location controls, compliance, reconciliation. The surface stays calm. The owner should not feel the machinery unless it is asking for a decision. The product leadership move was choosing that framing first, before a single screen was designed. ## What I did - Designed the owner experience around daily operational attention rather than raw configuration. What needs action, what is safe to act on, and what to trust are visible. Everything else is backstage. - Built a deterministic pricing model the business could inspect, explain, and trust. Pricing, discounts, commissions, and what-if scenarios route through the engine. The AI layer can explain a margin. It cannot compute one. - Created a multi-location architecture with role and access boundaries from day one, before the second location existed. - Added AI only where it supports decisions without transferring authority from the operator. The assistant answers plain-English questions about the business. It reads bounded functions and explains what the engine already computed. - Built Content Studio for in-portal marketing copy authoring: a brand-rules linter and PHI scrubber with a fully deterministic, zero-token core, then layered an AI reviewer, a brief-to-draft co-writer, and semantic library search on top as a second phase, without reworking any Phase 1 artifact. - Hardened the Zenoti PHI firewall by replacing heuristic regex detection with a catalog-membership allow-list at the persistence boundary. The shift moved the gate from "block known-bad patterns" to "require known-good catalog rows," which is the only approach that holds against novel field names or encodings. - Designed the employee onboarding hub with role-scoped, deny-by-construction access: staff see exactly one surface (their checklist), escalation is structurally impossible from the invite path, and the RLS deny sweep is proven by tests rather than asserted. - Added a profitability layer covering COGS variance, a margin watchlist with a redemption log, and a monthly P&L view, plus three read-only AI tools so the assistant can answer margin and variance questions without ever computing money itself. The decision that shaped the product was choosing which artifact would lead. The pricing engine was the impressive one: 69 services priced across tiers and roles, overrides written to immutable history, every number traceable. The obvious move was to make that grid the front door and let the owner admire the machinery. I killed that framing. The constraint that forced my hand was the PHI-free assumption I had drawn early: no appointment-level records, no guest names, aggregate operational signal only. That assumption stripped away the transactional detail a conventional admin surface would lead with, which left only aggregate signal to build from, which is exactly what a first-time owner needs on opening day. The privacy constraint did not fight the product. It produced the cockpit. So the grid moved backstage and the daily decision queue became the front door. The pricing engine was the impressive artifact. The cockpit was the right product. ## What changed The practice gained an operating surface before opening day. What began as a 72-hour sprint of 15 pages now covers roughly 17 surfaces, from the compliance sign-off the Medical Director stands behind to the month-end reconciliation that catches drift before it becomes a surprise. Scattered setup decisions became a visible system: what exists, what needs review, what is safe to act on, what to trust. - The owner replaced a folder of spreadsheets with one place she can stand behind: every service, membership, and expense lives in a system that prices and reconciles them the same way. - She opens the day to signal instead of silence. Below-cost services, unsigned attestations, expiring credentials, and reconciliation drift surface on their own instead of waiting to be discovered. - Questions that used to route to her now route to the system. Plain-English answers come from deterministic logic, not model-generated guesses, so the answer is the same one the books would give. - The second location is designed to be configuration, not a rebuild. The access boundaries shipped before it existed, so a new site slots in as configuration rather than forcing a migration. - Marketing copy now has a home inside the portal. Content Studio lints every draft against brand rules and scrubs PHI before anything saves, with all safety logic running as pure functions at zero token cost. The AI reviewer and co-writer layer on top for operators who want a second pass or a starting draft. - Staff onboarding moved from shared documents to a structured hub: each employee sees their own checklist, managers see a progress matrix and get flagged when anyone stalls, and scope of practice lives in the system rather than a PDF. - The profitability layer surfaces what the pricing engine implied but never showed: which services are losing money against their cost structure, how margins have moved month to month, and where discount usage is running against the cap. The AI reads the same numbers the books use and narrates them in plain English, nothing more. The design move was subtraction. The platform does a lot. The interface makes only the right things feel urgent. ## Reflection The through-line of this build is what I chose to leave out. The strongest story is not "15 pages in 72 hours," or the 17 surfaces it has since become. It is picking the right surfaces: the owner needed operational confidence, not software theater, and the product leadership sat in what I kept backstage. **A privacy constraint can be a design engine, not a tax.** The PHI-free assumption looked like a limit on what the product could show. Held early, it did the abstraction work for me. It forced aggregate signal, and aggregate signal is exactly what the cockpit needed to lead with. Next time I would reach for a hard constraint sooner rather than treat it as a cost to route around. **Some of that structure is a bet on a future that has not arrived.** I modeled multi-location access boundaries before the second location existed, guessing at edges I could not yet see. It shipped clean, but it was speculative architecture, and speculative architecture is a debt until a real second site proves the shape right. I would not draw those boundaries the same way again. ## Microsoft Source: https://mamjed.com/case-studies/microsoft # Microsoft ## TL;DR Partner onboarding for Windows 8 app experiences ran as a queue of one-off builds. I led the white-label initiative that turned it into a system the next partner could inherit, and onboarding got over 60% faster. As a product designer on the App Experience Team, I owned the design system, the motion, and the coded UI. It shipped and stuck across regions. ## Every partner started the same build from scratch Each partner felt like a new project: new asks, new assets, new coordination, new delay. The team absorbed that variance one engagement at a time. Onboarding depended on whichever designer happened to take the brief, and every new region paid the same setup cost again. The partners were not the problem. The process was. There was no shared backbone between partners, design, and engineering, so every custom deliverable started from scratch. ## What I did I stayed in design, motion, and coded UI rather than handing off at the mockup, so the gaps that open at every handoff closed instead. And I built repeatable patterns for regional requirements, brand assets, and engineering constraints, so a partner entered a system instead of a queue. I also ran a standing "Friday Tips" session inside the team, treating tooling and productivity knowledge as a shared asset so the gains compounded past any single engagement. But the move that made any of that possible was mapping the variance first. ### Partner variance was not infinite When I mapped what actually differed across engagements, the differences clustered into a small number of zones instead of a fresh set of one-offs each time. Most of what a partner needed was a shared shell every partner could inherit. A middle band of regional and brand options could be set rather than rebuilt. Only a thin layer was genuinely unique to each partner. That clustering let me draw the line in one place: make the middle zone configurable, hold the rest fixed, and stop treating every partner as a bespoke build. I could have kept scoping each engagement on its own, which would have felt safer and made no promises I might not keep. I chose the line instead, because the variance had shown me it was real and stable. The white-label zones came out of that mapping, not out of a template, and that is what moved the numbers. ## What changed - Onboarding time dropped by over 60%: the system carried the setup cost once, and every partner after that inherited it. - Speed stopped depending on which designer took the brief, and the workflow held across regions and engagements rather than fading after the first rollout. ## Reflection **The operating model moved the numbers, not the craft.** This was early proof of a habit, not a flagship. What paid off in a partner program was the workflow behind the deliverables: fix it once and the next partner inherits the benefit without anyone having to be a hero. It is the same instinct I have reached for in every systems role since. **The catch lives in how I drew the line.** I mapped the variance from the engagements I could see, then locked the zone boundaries early, and early boundaries are only as good as the sample they came from. If a later partner had needed something the configurable band could not express, the fixed shell would have fought me instead of flexing. On a longer program I would leave the seam between configurable and fixed easier to move, so a partner who revealed a case the first mapping missed could reshape the system rather than break it. ## Nibbble Source: https://mamjed.com/case-studies/nibbble # Nibbble ## TL;DR I founded Nibbble to give independent restaurants loyalty tools that fit how they actually operate. It is a live multi-tenant SaaS at app.nibbble.io: three role-specific portals, Square POS, Stripe billing, and a shared token layer, shipped 2026-04-22 and still running. A four-restaurant customer council validated it, two are in active beta, and it is pre-revenue by design. The strategy and design are mine end to end; Ahsan Amjed was my developer collaborator, with two others co-authoring commits. ## Three surfaces, not one role-gated app The easy version is a punch card with a login screen. Independent restaurants need a system that respects how they actually operate: the owner thinks about margin and repeat visits, staff need speed between orders, customers need a reason to return without homework. Started with the operating model, not the feature list. Nibbble is three products sharing one spine (owner, staff, customer), each with a different job on the same architecture. The harder challenge was keeping that system coherent while building primarily with AI agents. Without the right scaffolding, agents drift the system every week. ## What I did - Defined the product strategy for independent restaurants, not enterprise chains. The customer council of four restaurants validated decisions before they shipped. - Designed the core experience across owner, staff, and customer workflows. Three distinct surfaces with different primary actions: admin carries full configuration weight, staff is built around one high-frequency action during rush, customer is built for scan-and-go on a phone. - Led the production v1 build with Ahsan Amjed: Square POS integration, Stripe billing, multi-tenant Postgres with row-level security, and scheduled operational jobs. Launched 2026-04-22 at app.nibbble.io. - Created a shared design-token system so three portals feel like one product. Three distinct layouts, one coherent system. The fork that organized the design: one role-gated app, or three distinct surfaces. The cheap path was a single app that hid and revealed sections by role. It ships faster and shares one codebase. I rejected it because the three users are not the same person with different permissions. They have different anxieties and different tempos: the owner reasons about margin, the server needs one action during a rush, the customer wants scan-and-go with no homework. Role-gating would have forced a shared layout to serve all three and served none of them well. I shipped three distinct surfaces on one shared data model and token layer instead. The token layer is what pays down the cost of that decision: three layouts stay coherent without three separate design systems to maintain. ## The AI operating model that kept the system coherent Building primarily with AI agents drifts the system on a weekly clock. Left to the prompt alone, agents touch files they should not, let docs fall out of sync with code, and re-solve problems the codebase already solved. On a three-portal SaaS held together by a shared token layer, that drift is the difference between one coherent product and three that slowly diverge. So I put the safety in the runner, not the prompt, where an agent cannot argue its way past it. The auto-fix loop runs behind a hard path allow-list with a one-file-per-fix limit, a pre-push hook blocks any code change that skips a docs update, and CI requires a second-model review before anything merges. A codebase-memory graph lets agents query the system's structure before editing instead of grepping blind. That is governance with a receipt, not a slogan: the loop has landed 17 real commits against real failures, gated behind the allow-list rather than running as a demo. Speed did not become drift. ## Phase 2 shipped (2026-07) Four features that close the loop from product promise to measurable outcome, shipped via parallel worktree agents with Ahsan Amjed: - **White-label theming:** the customer portal now renders in the restaurant's brand colors and logo. Previously admin stored the brand data; Phase 2 applied it end-to-end. - **Baseline event log:** the north-star metric (repeat-visit lift) was previously unmeasurable because insights were snapshots. The event log makes it measurable for the first time. Cohort queries follow in the next phase. - **PWA home-screen layer:** diners authenticate via phone OTP and are mobile-first. The PWA manifest and service worker give the customer portal a native-feeling install path on the device they already have in their pocket. - **Notification delivery worker:** provider-pluggable (log default, Twilio env-gated). Enqueues on tier upgrade and punch-card completion with deduplication, so the notification loop is closed end-to-end rather than scaffolded and waiting. ## A security fix that became a class-level fix A punch-card bug (one restaurant could award loyalty value to another restaurant's customers) revealed a pattern across multiple staff-authenticated routes: each trusted a `restaurant_id` supplied in the request body rather than binding it to the caller's token. The routes were distinct but the vulnerability was the same shape. Fixed the routes, then widened the CI guard to cover every `verifyStaffAuth` route, not just the admin paths the structural test already watched. The class is now enforced at the CI layer, not just at the instance that surfaced the problem. ## What changed Nibbble moved from idea to live SaaS with a business model, a customer council, a production foundation, and a system for keeping product decisions coherent as it grows. - Production multi-tenant SaaS running at app.nibbble.io since 2026-04-22: three portals, Square POS, Stripe billing, multi-tenant data layer. - A customer council of four restaurants, two in active beta. Pre-revenue by design: at this stage the asset is operator-validated product decisions and production discipline, with revenue as the next milestone. - Phase 2 shipped: white-label theming, baseline event log (north-star now measurable), PWA, and notification delivery, wired end-to-end. - A documented multi-agent development stack where design, product, and engineering choices stayed connected through the build. ## Reflection **Encode the guardrail where it cannot be argued away.** Putting the allow-list, docs gate, and second-model review in the build system rather than the prompt is what let a small team move fast without the system fragmenting. A prompt can be forgotten or overridden; a runner cannot. That is the pattern I carry into any AI-native team I lead. **Stress-test the roadmap from the demand side, not just the engineering side.** Running a three-lens review (CEO, CTO, CDO seats, each grounded in a code audit before forming an opinion) surfaced that the roadmap was complete and well-sequenced for engineering while missing its entire demand side: no pilots, no positioning, no revenue motion. Only role-forcing exposed it. The output was a new "Prove it" phase inserted before product expansion: positioning one-pager, named pilots, and the baseline event log that makes lift measurable. The roadmap now carries all three lenses with provenance, ready for Ahsan and me to ratify owners and dates. Two things I would do differently: 1. I optimized for build discipline before I had demand proof. Nibbble is pre-revenue by design, but that framing can hide a real risk. I built runner-level rails, a docs-sync gate, and a token layer before a single restaurant paid. The next proof has to be usage, retention, revenue, or operator behavior, not more build velocity or a cleaner commit history. If I ran this again, I would spend the customer council on willingness-to-pay signals earlier, not just design validation. 2. Three surfaces was the right call, but I under-invested in the seam too long. The three-portal decision holds up. What I would change is timing: the shared token layer is the only thing keeping three layouts from becoming three design systems, and I treated it as infrastructure to harden later rather than the load-bearing decision it was. Running this with a Director of Product Design team, a design-systems lead would own that token layer from day one. Nibbble is what it looks like to turn ambiguity into a shipped product without letting it fragment. ## Salesforce Essentials Source: https://mamjed.com/case-studies/salesforce-essentials # Salesforce Essentials ## TL;DR Packaging enterprise software for a five-person team is mostly deciding what to leave out. As senior product designer for SMB and Salesforce Essentials from April 2017 to June 2020, I owned the packaging: what to show, hide, automate, and how much Einstein AI a five-person team met on day one. Engineering owned buildability, research owned the SMB study base, and I co-led the cross-company learning workstream. Essentials shipped and grew into a standalone business unit. Salesforce had deep enterprise capability, but SMB customers needed a first experience simple enough to grow into rather than bounce off. That framing is where the work started. ## The problem was not simplifying Salesforce "Make it simple" is easy to say and usually lazy. The harder question: simple for whom, at what moment, and at what cost to the user's future? Small businesses did not need a watered-down Salesforce. They needed the right first hour: what to do, what to ignore, how to get value before becoming CRM experts. The work became a series of packaging decisions. What to ship in Essentials. What to leave out. What to explain. What to automate. Where AI could remove setup instead of adding a feature to learn. The surface decisions mattered, but the upstream question mattered more: which Einstein capabilities actually helped a five-person team make better decisions faster? Engineering asked which ones were buildable. Design kept asking the different question. ## What I did - Designed the SMB surface to meet users earlier in their maturity curve, packaging the experience around the few jobs a five-person team needed to finish on day one instead of the full enterprise feature set. - I drew the line on which Einstein capabilities shipped in Essentials. Data capture and email parsing went in. Predictive lead scoring stayed out. Engineering could build scoring, and it demoed well, so the pull to include it was real. I rejected it because scoring asks a five-person team to trust a black box before they trust the CRM itself. It hides its reasoning at exactly the moment a new SMB user is deciding whether the product is worth the effort. Data capture and parsing earn trust by removing setup the user can see; scoring spends trust the product has not yet earned. So the rule I held was: AI that removes work ships now, AI that asks for faith waits until the customer has a reason to give it. - Rebuilt onboarding as a product with its own users, its own success states, and its own design problem, replacing a set of per-feature modals with one coherent adoption journey. - Co-led the cross-company workstream to align learning experiences across the Salesforce ecosystem around that frame, moving from fragmented per-feature tutorials toward coherent cross-product journeys. The Einstein line I held inside Essentials went on to shape how the company thought about AI for SMB beyond this product. The move underneath all of it was to reduce cognitive load without reducing ambition. The product could still grow with the customer because the architecture was never capped at simple. ## Reflection 1. Packaging is subtraction, and subtraction is the hard part. The strongest version of this work is not "I made Salesforce simpler." It is deciding how much enterprise power a small business should meet at each stage of maturity, and defending that call inside a company that wants to ship everything. Saying "this capability is real but it is not right for this user at this moment" is a harder judgment than it sounds, and it is the skill I carry into every role since. 2. What "packaged" cost us: speed. Rebuilding onboarding as a single coherent journey instead of shipping per-feature modals was the right call, but it was slower to get to first release, because a coherent journey has to agree with itself end to end while a modal only has to explain one screen. Individual teams could have shipped their own tooltips in a fraction of the time. If I ran it again I would ship a thinner first cut of the journey sooner and widen it in the open, rather than holding for coherence I could have earned incrementally. 3. Holding back predictive scoring was correct for launch, but I never closed the loop on when it should ship. I set the gate ("wait until the customer has a reason to trust it") without defining the signal that would open it. A cleaner version of the decision names the trigger, not just the veto. ## SalesforceIQ Source: https://mamjed.com/case-studies/salesforce-iq # SalesforceIQ ## TL;DR The Contact Gallery redesign lifted daily active users 40% and duplicates merged 34%. SalesforceIQ's back-end knew things about relationships its front-end could not turn into action. As Principal Product Designer, I carried the contact intelligence and data-quality surface from research through entity modeling to interaction design. I redesigned Contact Gallery around action, trust, and correction with multi-merge and a master contact model. I worked weekly with support, product, marketing, and engineering, and later fed early Einstein UX thinking in a consulting capacity. It shipped between December 2015 and April 2017. ## The intelligence was there. Users could not act on it. SalesforceIQ was early to the AI CRM story. The back-end captured relationships traditional CRM could not see. But intelligence in the back-end does not matter if the front-end makes users suspicious, confused, or tired. Users still had to answer the same questions on every visit: Is this the right contact? Can I trust this record? What should I do next? What happens if I merge? Three long-standing customer issues had been open for months. The product was not translating what the AI knew into anything a sales user could act on without re-learning the app. I worked weekly with support, product, marketing, and engineering so prioritization stayed honest about customer sentiment, then took the contact intelligence and data-quality surface directly. ## What I did - Researched pain points with AEs, customer success, and customers to anchor the redesign in actual friction, not assumed friction. - Shipped fixes to three long-standing customer issues in the first month before opening a broader redesign conversation. The credibility bought license for the harder work. - Reframed Contact Gallery from a display surface into a workflow surface: designed for what users needed to do (understand, trust, resolve), not for how to show more data. - Introduced multi-merge and a master contact model, which meant rethinking the contact entity, not just adding a button. The merge work is where the real fork showed up. Customers knew they had duplicates, so the obvious read was that merge was too hard to find or too tedious to run, and the fix was a faster, bulk-merge button. When I watched how people actually behaved, the friction looked less like effort and more like fear: merge collapsed several records into one in a single irreversible move, and users would not spend trust they could not get back. So I did not optimize the old action. I made merge composable and previewable. The user picks the master, chooses the default photo, and edits the assembled contact while every name, email, phone, and handle stays visible, all before anything is committed. Reframing merge from one irreversible action into a sequence of reversible, inspectable decisions is what moved the number, not a faster button. Two things sat next to the merge work. After the PredictionIO acquisition I scoped the Einstein contribution honestly as adjacent, pre-product groundwork and worked it in a consulting capacity: how a confidence signal, a source, and a correction path should read when an AI surfaces an insight a salesperson has to act on. I did not own an Einstein deliverable, and I claim no metric for it. Alongside that, I mentored the first AI-focused designer on the team. ## What changed - Daily active users increased 40% after the Contact Gallery redesign. - Duplicates merged increased 34% after multi-merge and master contact model shipped. - Three long-standing, high-frequency customer issues closed inside the first month. - Early Einstein consulting on confidence, source transparency, and correction loops fed forward as pre-product groundwork, before that work had a named owner. Trust improved because users got clearer control over machine-captured data. AI earns trust by making the next action clearer, not by making the data denser. ## Reflection **Trust was a control problem, not a data problem.** The intelligence layer was already ahead of the experience, so the lift never came from surfacing more of it. It came from giving users reversible, inspectable control over machine-captured data. Closing three long-standing customer issues in the first month is what bought the license to rethink the contact entity at all. Had I opened with the redesign, I would have been arguing for a model change from zero standing. Credibility first, then the control work, then the number. In that order it held. **Where I would work differently.** I ran the future-state visioning in parallel with the near-term product work instead of letting a clearer future thesis inform the near-term calls. That cost the redesign a through-line: the work landed as a set of good decisions rather than one argument built backward from where the product was headed. Next time I would settle the thesis first and let it discipline the near-term calls. This work is a bridge between earlier CRM design and today's AI product work: the same problem of making machine intelligence something a person can trust and act on. ## Samsung Ads Source: https://mamjed.com/case-studies/samsung-ads # Samsung Ads ## TL;DR I helped start Samsung Ads and ran design across it. Mine: the interaction work, the front-end engineering, and the data-visualization layer buyers read. Data science owned the signal models underneath; I partnered with them on what to surface. The proof: a tiny team became a standing Samsung business unit that turned a profit in year one, and the way I turned aggregated viewership into readable patterns became a filed patent. ## Building an ad category that did not exist yet The real problem was creating a new business inside a large company. You design product surfaces, but also trust with executives, workflows with data science, standards with partners, and enough operating structure for a small team to survive inside a much larger organism. Samsung had the hardware footprint and first-party smart TV data at a scale no independent network could match, and no ad business to monetize any of it. The question was not whether smart TV advertising could be a business. It was whether a new team could move fast enough inside Samsung to prove the category before the window closed. Small team, big ambition, high ambiguity. We sat close to product and engineering because that was the only way to move fast enough. It worked: the category held, and the small team became a standing Samsung business unit. ## What I did - Co-founded what became Samsung Ads as one of four founding members. Charged with building the design and front-end side: the advertiser platform, the data visualization layer, the standards posture with IAB, and the team behind all three. - Led design and front-end engineering for the early platform. Held a quality bar that matched advertiser trust expectations from day one rather than deferring polish to a later phase. - Contributed to IAB working groups on smart TV and large-display advertising standards rather than waiting for external rules to land. That gave Samsung Ads early influence over the category and produced tighter design constraints for the platform. - Mentored product managers, supervised a satellite team, and briefly served as general manager during the scaling phase. - Aligned product, platform, and design standards across US and Korea collaboration throughout the build. **The key discovery: buyers act on patterns, not raw signal.** The first cut of the dashboard surfaced every viewing signal Samsung's TVs produced. Advertisers froze. More data did not read as more insight; it read as noise, and a buyer who cannot decide does not buy. Working with data science, I found that buyers acted on a small set of recurring viewing patterns, not the full firehose. So I designed the platform to surface those patterns and hide the rest. That framing, turning aggregated viewership into readable patterns, became the patent I filed (`samsung-patent`). ## Outcomes **A founding experiment became a standing business.** The team went from four founders to a Samsung business unit, the pattern-first platform shipped as the advertiser daily driver, and the viewership-to-patterns approach became a filed patent. The design language I set for it carried into the broader Samsung TV experience and stayed there. ## Reflection 1. Speed inside a large company is never just speed. It is negotiation, alliances, and enough proof to keep the corporate antibodies away. The dashboard was as much a political instrument as a product: a small team survives by shipping something executives cannot argue with. 2. More data is not more insight. The first cut drowned advertisers in signal; the platform only worked once it surfaced the few patterns buyers actually acted on. I now reach for the smallest legible view first, not the most complete one. 3. What I would do differently: engage standards bodies earlier. We contributed to IAB working groups, but reactively, responding to agendas others set. Helping set that agenda from the start would have produced cleaner product constraints and a stronger competitive position from day one. ## Simple Cortex Source: https://mamjed.com/case-studies/simple-cortex # Simple Cortex ## TL;DR Most AI consulting starts with prompts and ends with drift. I built Simple Cortex to sell the missing piece: governance, designed into the operating layer instead of buried in a prompt. It is a live full-service firm, web and brand and strategy at the core, run since 2025 out of Northern Virginia, with active client engagements. AI is the operating edge, not the headline identity. As founder and principal I built the operating model end to end: the agent contract, the routing layer, and the rules every contributor reads. That system is what the firm sells and what carries each engagement I lead. ## AI was never the constraint. The operating model was. Everyone has access now. That is not the hard part. The hard part is knowing what work goes to which model, what it should cost, what voice it should use, what tools it can touch, and when a human must review. Without answers to those questions, AI fluency is just speed without quality. Engagements drift. Outputs lose consistency. Cost is invisible until it is a problem. Simple Cortex treats AI like an operating model, not a shortcut. Work is routed. Voice is governed. Instructions are inspectable. Cost is visible. Judgment is designed into the system instead of buried in a prompt. ## What I did - Built the practice around governed delivery, not one-off prompting: agent personas, execution loops, tool surfaces, and review checkpoints all defined as reviewable artifacts committed to version control. - Designed a local-first routing layer. The OpenClaw Router runs locally and classifies work across six semantic routes (code, research, organization, conversation, visual, architect), sending most of it to local models and reserving cloud fallback for the calls that need the most judgment. - Named the registry of seven models the router chooses among, so routing is concrete, not abstract: a local classifier (llama3.2:3b) triages, local models handle companion (qwen3.5:9b), engineering (qwen3-coder-next), analysis (openclaw-qwen35-a3b-think), and vision (qwen3-vl) work, and two cloud models sit at the top of the ladder: a senior model (claude-sonnet-4-6) and an architect model (claude-opus-4-6) for the hardest calls. - Created production agent infrastructure. Three agents in production (a CEO, a Founding Engineer, and a Researcher), each defined by a four-file contract, so persona, instructions, tools, and execution rhythm are independently reviewable. - Codified voice, tool access, and decision rules as living documents that any contributor, including AI contributors, reads before touching client work. - Applied the system across active client engagements in brand, product, automation, and AI-enabled operating systems. The decision worth defending is how agents ask for a model. I made agents call the router by intent rather than name a model inline. The alternative, letting each agent hardcode its own model ID, reads simpler at first: the model an agent uses is right there in its definition. I rejected it. Binding an agent to a model ID means a swap, whether for cost, quality, or a provider outage, becomes an edit across every agent that named that model, and the cost of each call disappears into the agent instead of surfacing at the routing layer. Calling by intent puts model selection, fallback, and cost accounting in one place the operator can inspect. Portability and one line of routing change beat vendor loyalty and a sweep across prompts. ## Reflection Simple Cortex supports rather than competes with the product work. It is the operating discipline behind the current ventures, including the Kintsu Medspa engagement. A practice that runs lean on governed local inference can sell that discipline back to its clients, not just the output. The forward move is pulling the persona and voice files into a shared style guide so every new agent inherits the rules by default, the same way a design system enforces token discipline across components. Two things I would tell anyone building the same thing: 1. Governance is the product, not the overhead. The instinct on a fast engagement is to skip the contract and just prompt. I built the four-file contract and the routing layer first anyway, because the alternative is speed that quietly loses voice and blows past cost. The discipline is what a client is actually buying. 2. Local-first is a real tradeoff, not a free win. Routing most work to local models keeps cost visible and low and keeps client data off third-party servers by default. It also means I own the latency and the reliability that a cloud vendor would otherwise absorb, and I keep two cloud models on the ladder precisely because the cheapest local path is not always the right one. Choosing where judgment belongs is the design work; pretending local is always better would be a lie. ## Sitetracker Source: https://mamjed.com/case-studies/sitetracker # Sitetracker ## TL;DR From June 2021 to March 2023, I was Sitetracker's Head of Design and built the design function from nothing: the hiring bar, the career ladder, the critique culture, and a 4-person international team held to the same standards as the US group. Development ran closer to waterfall than to customer value, so I made Jobs to Be Done the shared decision language across design, product, engineering, and QA. The Head of Product credited the work with accelerating the whole product organization. ## Design had no system to hire into Hiring designers into a broken system creates frustrated designers, not a design function. Sitetracker needed the conditions for design to matter first: a hiring bar, a career ladder, critique, mentorship, research language, and a product process that made customer value visible before work was committed. The company had built a real business without a real design function. Design sat downstream. Reviews happened late. Product, engineering, and QA had no shared language for the decisions they were making together. The real constraint was how the company made product decisions, not the size of its org chart. ## What I did - Assessed product maturity before defining the org shape. The skill matrix mapped the capabilities the business actually needed rather than a generic design org chart borrowed from somewhere else. - Wrote the design career ladder and mentorship framework before scaling headcount. New designers walked into a system that already told them what growth looked like. - Introduced JTBD across design, product, engineering, and QA through workshops and coaching. The framework stuck because it lived in how people argued about product decisions, not in a document that gathered dust. - Held a peer seat in weekly leadership reviews of features and projects. Authority was earned, not inherited. - Built and trained a product and design team internationally for global expansion. That team operated as a peer node from the start, not a delivery arm. I sequenced the system before the people. The career ladder, mentorship framework, and skill-matrix hiring loop were all in place before the team grew past the first few hires, so every new designer landed inside a function that already worked instead of one being improvised around them. The matrix only worked as a hiring tool after I stopped scoring credentials. The obvious move was to write job specs and filter for seniority and pedigree, the way most design orgs staff up. I built the first version that way and it reproduced the exact gut-call bias I was trying to remove: strong resumes advanced, and the real capability gap stayed open. So I inverted it. I scored candidates on demonstrated capability against the specific gaps the matrix named, not on where they had worked. That version cut the bias, and it made every hiring decision legible to product and engineering leadership, who could then co-validate a hire instead of deferring to my taste. Credentials describe who a candidate has been. The matrix describes who the function needs next, and only the second question builds the right team. ## What changed The change I care about is not that design got better screens out the door. It is that JTBD stopped being a design research method and became the language product, engineering, and QA all argued in, so decisions got made against customer value instead of opinion. Design moved from a late-stage review gate to a co-decision partner at the start of each initiative, and that shift is what accelerated the product organization. Late-stage rework fell too, because the hard calls happened together and upfront instead of surfacing in a review at the end. Two people who worked with me name the same thing from different seats. Bailee Warsing, a designer on the team, wrote that I "collaborated frequently with Product Management and Engineering leadership to build a robust design playbook." That is the inside-the-team view of the cross-functional standard the Head of Product describes from the leadership side. The international group stood up on that same standard, training, and culture, so global expansion never carried a quality cliff at the boundary. And the whole operating system kept running after my tenure. That is the artifact that mattered: not the screens, but the machinery that made better screens possible without me in the room. ## Reflection 1. **I would push JTBD outward earlier.** I ran the workshops and coaching, but let the design function consolidate for a stretch before extending the methodology to product and engineering. Starting both threads at once would have compressed the culture-change curve and given the rest of the org more reps with the framework before it had to carry real product decisions. 2. **Building the international team as peers cost me speed early, and I would pay it again.** I stood the group up on the same hiring bar, ladder, and critique culture as the US team instead of as a faster, cheaper delivery arm. That was slower to staff and slower to ramp, and for a stretch a delivery-arm model would have shipped more, sooner. But a peer node holds the standard when no one is watching, and a delivery arm does not. Wrong call for the next two quarters, right call for a function that outlasts my tenure. # Writing ## The 3 Startups You'll Find in Large Companies Source: https://mamjed.com/writing/three-startups-inside-large-companies # Three Startups Inside Large Companies Three times in my career I have helped build something new inside a company much larger than the team building it. Samsung Ads inside Samsung. SalesforceIQ inside Salesforce. Salesforce Essentials inside Salesforce again. Each time the work looked like a startup from the inside and looked like a feature from the outside. Each time the same forces showed up. Each time the work taught me something specific about how to lead at the seam where a small team is trying to build a real business inside a much larger one. This is not a celebration of intrapreneurship. It is a description of the pattern. ## The three anchors **Samsung Ads.** Zero-to-one new business inside a global hardware company. Samsung had massive smart TV reach and no ad business. I co-founded the initiative as part of a four-person team and led design and front-end engineering. We built the advertiser platform, the data visualization layer, the standards posture, and the team itself. In its first year the business generated $20M in profit. We filed a patent for AI-powered visualization of global TV viewership behavior and contributed to IAB standards for smart TV advertising. This was building a category, not a feature. **SalesforceIQ.** An acquired AI-assisted CRM absorbed into Salesforce and re-shipped. The back-end was ahead of the front-end. Users could not trust or act on the intelligence the product surfaced. I redesigned the Contact Gallery, introduced a master contact model and multi-merge capability, and consulted on early UX foundations for what became Einstein. The Contact Gallery redesign increased daily active users by 40% and duplicates merged by 34%. Different startup, different stage: a product that already existed had to find its place inside a platform that already existed. **Salesforce Essentials.** SMB packaging inside an enterprise platform. Salesforce was too complex for small businesses to adopt. Essentials was the answer. A focused entry point for teams of ten or fewer that preserved a path into the broader ecosystem. The work was figuring out which Einstein capabilities belonged in Essentials, how onboarding should actually work, and how to align customer learning experiences across the company. A startup again, this time defined by what to take out, not what to add. ## What is the same across all three > The work is real. The container is borrowed. A few patterns showed up every time. **Oxygen is the scarce resource.** Not money. Not headcount. Oxygen. Executive attention, calendar time, a clear lane that other teams cannot drift into. Inside a large company there are always teams that have been there longer, have more relationships, and want the surface you are trying to build on. You spend a non-trivial part of your week defending the existence of the work, not doing the work. **Credibility is earned in the first month, not the first quarter.** The team you are inside of has already decided whether you can build. Show up, close a known problem fast, and the rest of the work gets easier. At SalesforceIQ that meant resolving three long-standing customer issues in the first month before touching the Contact Gallery redesign. At Samsung Ads it meant shipping the early platform at production quality even when the temptation was to ship rough and polish later. License is earned by shipping, not by framing. **You are accountable to two different customers at once.** The external customer, who has to find value. The internal customer, who has to find a reason to keep funding the work. The product has to land with both. A startup outside the building only has the first one. A team inside the building forgets about the second one at its peril. **You move between levels constantly.** A specific UI pattern in the morning. A product principle by lunch. A standards conversation in the afternoon. A staffing decision before you log off. The luxury of one altitude at a time does not exist inside an internal startup. The job is to be coherent across all of them. **You inherit constraints you did not pick.** A brand. A platform. A sales channel. A reporting line. Some of these constraints are gifts. Some of them are taxes. The work is figuring out which is which fast, then designing inside the gifts and around the taxes without complaining about either. ## What is different from a clean startup > A startup outside the building decides what to build. A startup inside the building decides what to build and what to defend. **You do not get to pick the cap table.** Your investors are your executives, and they are also your competitors for attention with other internal teams. You cannot fire them. You cannot dilute them. The political layer is not a distraction from the work; it is part of the work. **The exit is not an exit.** A clean startup is built to be acquired or to go public. An internal startup is built to be absorbed, integrated, or quietly wound down. The endgame is integration into the parent, and the integration almost always changes the product. If you do not design for that future you will get a worse version of it imposed on you. **Speed is bounded by the slowest dependency you cannot replace.** Legal. Security. Brand. Procurement. A clean startup can route around almost any internal bottleneck by hiring or by changing tools. An internal startup has to negotiate with the existing organism. The leadership skill is knowing which fights are worth picking, which are worth losing on purpose, and which are worth routing around without making it personal. **You do not get to pretend the parent does not exist.** Samsung Ads had to fit into Samsung TV. SalesforceIQ had to fit into the Salesforce platform. Essentials had to fit into the enterprise pricing and packaging story. The pretend-we-are-a-startup posture works for about a quarter, and then the parent shows up. Designing for the eventual handshake from day one is cheaper than retrofitting it. **Talent dynamics are inverted.** A clean startup hires people who want autonomy and equity. An internal startup hires people who want autonomy without leaving a stable company, which is a smaller pool and a more specific personality. You spend more time finding the right small group of people than you would in either pure context. ## What this builds in a leader A leader who has built three startups inside large companies has learned how to make new things happen inside organizations that are not naturally built for new things. That is not a generalist skill. It is a specific one. It looks like: - Building credibility with skeptical executive sponsors quickly, without performing certainty you do not have. - Holding a quality bar at startup speed, because trust inside the parent and trust from the customer are won by the same thing: work that looks like it knows what it is doing. - Designing for the eventual handshake with the parent organization from day one, so the integration is something the team shaped rather than something done to it. - Moving fluidly between a UI decision, a packaging decision, a standards decision, and a staffing decision in the same week. - Knowing the difference between a constraint and a gift, and not wasting energy fighting the wrong one. This is the skill that makes a senior design and product leader useful in an inflection-point company. The kind of company past pure-startup but not yet a steady-state machine. A company with a real business and a new bet inside it. A company that needs someone who has done this before, can read the room, and can ship the work. ## Close I did not set out to build a career of internal startups. The pattern accumulated. Looking back, the through-line is clear: I am most useful at the seam where a small team is trying to build something real inside a much larger thing, and where design, product, business, and politics all have to be held in the same hand. Three times now I have done that work and watched what it produced. The fourth time, whenever it shows up, will look like the same job in a new container. ## How I Branded Muslim Youth Source: https://mamjed.com/writing/branded-muslim-youth # Branded Muslim Youth ## The folding table The first time I saw it work, I was standing in a parking lot behind a community center, watching a hundred young men in matching tees pack hot meals into foil trays. The line was quiet. Nobody was performing. There was a brand on the shirts and a system behind the table, and the two were doing the same job: telling everyone in the room what we were here to do, and what we were not. That parking lot is the reason I write about brand the way I do. A brand is not a logo. A brand is a promise made legible enough that strangers can keep it without being told. ## What AMYA is The Ahmadiyya Muslim Youth Association is the youth arm of the Ahmadiyya Muslim Community in the United States. It runs as a national volunteer organization: chapters across dozens of states, a leadership cadence that turns over every few years by design, and a mandate to translate the faith's first principles (loyalty, service, knowledge, integrity, dignity) into work in the communities we live in. I serve as VP of Design and Innovation. The shorthand is simple: we are young Muslims who show up, in our cities, for the work that needs doing. ## What changed when we treated it like a brand and an operating system For a long time the work was real but the picture of the work was scattered. Every chapter ran its own service day. Every flyer looked different. Every social account told a different story in a different voice. Outside the community, people did not know what AMYA was, even if they had eaten a meal we packed. A few decisions reshaped that. First, we wrote the brand down. Mission, audience, voice, and the things we would refuse. We named the audience as two: the young men inside the organization, and the neighbors we serve. That second audience was a forcing function. It pulled the work toward outcomes a stranger could see, not internal rituals only we understood. Second, we set an identity system that traveled. One wordmark, one type pairing, one palette, and templates that a regional officer could populate without asking permission. The point was not control. The point was that a flyer made in Houston should be recognizable to someone scrolling past it in Boston, so that what we kept saying actually accumulated. Third, we built an operating cadence on top. A national campaign calendar. Quarterly themes the chapters could plug into. A shared playbook for the things we ran every year, like Ramadan food drives and Thanksgiving relief days, so a new local leader inherited a working machine instead of a blank page. Volunteer organizations live and die on continuity. The cadence was the continuity. Fourth, we narrowed. We said no to a long list of well-meaning side projects so we could say yes, fully and visibly, to hunger relief, blood drives, and civic service. Focus, in a volunteer org, is the single hardest decision and the most generous one. It tells your people where their time will actually count. ## What it added up to Over the years our chapters have raised and donated over $100,000 toward hunger relief, and helped feed over 700,000 hungry Americans through coordinated food drives and community kitchens. Those are the numbers. The thing the numbers point at is harder to count. A high schooler in our Detroit chapter now knows how to plan a campaign, brief a designer, brief a press contact, run a service day, and stand in front of a county official to explain who we are. That is the dividend of running a community organization as if it were a serious operating system. The next generation inherits competence, not chaos. ## What this transfers to product work The discipline is portable. The same things that made AMYA legible to a neighbor at a food line make a product legible to a customer in onboarding. A clear promise. A small number of decisions held with conviction. Templates a team can ship without re-deciding the basics. An operating cadence that turns one good day into a hundred consistent ones. When I sit down to shape a product org, I am running the same play. Write down what we are and what we refuse. Build the system that lets a new hire ship to the standard in their second week. Pick the cadence the team can hold. Narrow until the work is unmistakable. I am a faith-rooted operator. The same patience I bring to a critique room, I bring to community work. Both rooms reward the same craft. ## Close The folding table is still out there. So is the brand on the shirts and the system behind it. Most weekends, somewhere, a chapter is running the play. That is what the work was for. Not the logo, not the campaign, not the deck. The play, run again, by the next set of hands. ## Voice as a Lint Rule Source: https://mamjed.com/writing/voice-as-a-lint-rule # Voice as a Lint Rule I was sitting across from the founders of Kintsu, a new medspa, talking about how their brand should sound. Not adjectives. Sentences a real person would actually say to a guest at the door. We argued over a few words. One of them was "journey". They had heard it in every medspa pitch they had ever sat through. By the time we got up from the table, we had five voice pillars and a stack of SAY THIS / NOT THIS pairs. "Journey" was on the wrong side of the line. That is the point. A banned-words list is to brand voice what a typecheck is to code. It will not catch every problem. It catches the dumb ones you cannot afford to ship, and it runs every time without getting tired. ## The voice document had no teeth For a long time, a brand voice lived in a Google Doc or a Figma PDF. A designer opened it the first week of onboarding. A senior skimmed it before a launch. The rest of the time it sat on a shelf, and the brand drifted toward whatever felt fine on a Wednesday afternoon. Nothing in the workflow forced anyone to consult the doc. Nothing in review flagged a regression. Voice drift was a slow leak you noticed at the end of a quarter, when the homepage suddenly read like every other medspa on the internet. The drift is not loud. It is boring. That is the part people miss. ## The team is no longer just humans Half the copy on a modern project is drafted by a designer with Claude open, an engineer using Cursor, or a contractor with ChatGPT in another tab. The agents are fast and capable. They are also untrained on your brand. So the voice has to live where the agents read. > Voice walked with the client. Voice committed to the repo. Voice enforced at PR time. ## The cheapest possible voice gate The Kintsu prohibited-words list bans "journey", "holistic", "glow up", "self-care", "state-of-the-art", and "we believe". The point is not pedantry. The point is that a one-line check runs in code review and never gives the writer the benefit of the doubt at 4:55pm on a Friday. It does not care whether the writer is junior, senior, human, or a model. The line is the line. This is not a new idea. Engineers have encoded judgment into lint rules and CI for two decades. The new move is to put brand judgment into the same machinery, before the team can ship around it. ## The deliverable changed Six long-form Kintsu strategy documents, around 2,745 lines, now live in the repo. Positioning. Voice. Imagery. Naming. Tone for in-app moments. The deliverable used to be a deck. Now it is a set of files the team and the agents both read, in the place where the work happens. That changes what a senior brand or design hire is asked to do. The job is no longer to write the guide and hand it off. It is to encode the taste so the system enforces it. ## What this means for a design org Voice review stops being a copy review job. It becomes a contract a director writes once and the team enforces in code review, the same way they enforce a typecheck. A junior can hold the line because the line is in the file, not in their head. Taste still gets set in a room, in person if it can be. What changes is what happens after the room. The taste gets translated into rules a system can run. Banned words. SAY THIS / NOT THIS. A canonical facts file. A short markdown contract at the top of the repo. ## What the job actually is now Set the voice in person. Codify it. Put it where the agents read. Enforce it at PR time, alongside engineers enforcing types and security enforcing scopes. Skip that last step and the voice document goes back on the shelf, and the brand drifts. I have watched it happen. The rule in the repo is the thing that stops it. ## Credibility Before Vision Source: https://mamjed.com/writing/credibility-before-vision # Credibility Before Vision A new design executive shows up on a Tuesday with a deck. Twelve slides. A North Star. A maturity model. A six-month roadmap. By Friday the engineering lead has stopped reading the Slack thread. By the end of the month the team has quietly decided the new person is "all strategy." I have watched this happen more times than I want to count. It is not a smarts problem. It is a sequencing problem. New leaders lead with vision when they should be leading with a closed problem. ## What the team is actually deciding in week one Inside the team you just joined, a question is already being answered before you write a single line of your operating plan. Can this person build, or do they just talk about building. If you spend your first three weeks listening, mapping, and producing a deck, the answer the team writes down is "talks." Whatever the deck says is then read through that filter. The same strategy in the same words lands differently depending on what the team has already concluded about you. > Vision without credibility is a tax on your future self. You can still get the strategy approved. You will pay for it in slower execution and longer cycles to ship anything real. The team treats your direction as something to wait out rather than something to run with. That tax compounds quietly for months. ## License is earned by shipping, not by framing At SalesforceIQ, before I touched the Contact Gallery redesign, I closed three long-standing customer issues in my first month. Nothing in those three fixes was strategic. They were specific, named, in-the-backlog problems the team had not been able to clear. Once they were closed, the room changed. The Contact Gallery redesign, which eventually drove a 40% lift in daily active users and a 34% lift in duplicates merged, was scoped and reviewed by a team that had already decided I could ship. The redesign was the vision. The three fixes were the license to attempt it. At Axios HQ, the 90-day re-architecture of the IA worked the same way. The IA work was scoped fast and shipped fast so the larger play, doubling AI usage across the platform, had ground to stand on. The sequence mattered. Land something the team can point at, and the bigger swing gets heard as a real plan rather than a pitch. ## Different audiences read credibility differently Engineering reads credibility as "this person makes good calls and does not waste our cycles." They watch how you handle the first hard technical tradeoff. Do you decide cleanly. Do you understand enough of the stack to argue on the merits. Product reads it as "this person can hold a line on what matters and cut what does not." They watch how you scope. A leader who reduces scope to something shippable is read very differently from one who expands scope to something aspirational. Executives read it as "this person closes loops." Did the thing they said they would do in the first month actually happen. Was it visible. They are pattern-matching against every other leader they have hired, and the pattern they trust is closure. These reads happen in parallel, and each audience compares notes with the other two. One closed loop all three can see is worth more than three separate gestures aimed at each of them. ## The trap: confusing motion with credibility The reflex of a new leader is to introduce process. A new ritual. A new template. A renamed standup. These changes are easy to make in the first month because they need no one's permission. They look like progress. They are not credibility. Process changes signal that you are in the chair. They do not signal that you can build. A team that watches a new exec reshape the roadmap template before shipping anything reads that as motion. The question they are answering, can this person build, stays open. A closed problem is the only currency that closes the question. Three customer issues. One IA re-architecture. One redesign that moved a number. The team has to be able to say, in one sentence, what you actually did. ## Vision lands after credibility, not before I am not against vision. I am against the order. After a closed problem, the strategy work gets heard differently. The room leans in. The skeptical engineer asks how, not whether. The product partner brings their concerns into the plan rather than reserving them for the hallway after. The executive sponsor stops asking for weekly proof points and starts giving you cover. The strategy you arrive with on day one and the strategy you propose on day forty-five can be word-for-word identical. The reception will not be. That is the whole point. ## What this means for the hiring panel The question worth asking a candidate is not "what is your design philosophy." It is "walk me through the first month of your last role in specifics." A leader who has run this play can name the problem they closed, who was watching, and what changed in the room after. A leader who has not will reach for a framework. The first month decides the next year. A leader who knows that is a leader who has done this before. ## Useful, Not Theatrical Source: https://mamjed.com/writing/useful-not-theatrical # Useful, Not Theatrical At Axios HQ, the AI work that moved usage was the work that did not look like AI. No wand icon. No banner renamed "AI-powered". Better defaults. Suggestions placed where people hesitated. The right next step pre-filled. In the same window, I led an eight-person product and design org through a 90-day re-architecture of the information architecture. AI usage doubled across the platform. That number did not come from a banner. It came from features that did not announce themselves as AI. I have shipped AI inside three companies and two ventures. The pattern keeps holding. The work that moves usage is invisible. The work that announces itself is theater. ## The model is an input. Trust is the product. SalesforceIQ was an AI-assisted CRM before that phrase had a marketing budget. The back end was real. The intelligence was real. Users could not act on it. The product surfaced relationship signals the model was confident about, and the user was not. So they ignored them. That is the part most teams miss. Until a user can trust an inference enough to act on it, the inference does not exist. The model is an input to the design problem, not the answer to it. When we redesigned the Contact Gallery and added a master contact model with multi-merge, daily active users went up 40% and merges went up 34%. Same intelligence underneath. We changed where it surfaced and how a user could verify it in one look. The path to trust got shorter. ## Useful AI is constrained AI. The constraint is the feature. A constraint that lives where the work lives gets followed. One that lives in a separate Confluence page gets violated within a quarter. Useful AI inside regulated work is not the AI that does more. It is the AI that has been clearly told what not to do, in the same file as the rest of its job. ## Show the work, including the cost. Simple Cortex runs an open-source router called OpenClaw. Semantic routing across six routes: code, research, organization, conversation, visual, architect. The thing operators notice first is not the routing. It is that `ocr costs` is a daily-driver command. The cost meter sits next to the other tools, in the terminal where the work happens. That choice treats the operator as an adult who makes economic tradeoffs all day. Hiding cost behind a dashboard treats them as a user who should not worry about the meter. Both designs have a model of the user. Only one is correct for serious work. Trust is the same problem here as it was at SalesforceIQ. Operators believe the system more when they can see what it is doing and what it costs to run. Legibility is the feature. ## The theatrical failure mode The recognizable failure mode is the AI that announces itself. A wand icon. A "Powered by AI" pill. A two-second animation to show a model thinking. This always loses to a quiet feature that just works. Theatrical AI sets an expectation the model cannot keep. It puts the trust burden in the wrong place. The wand says it is smart. The user decides if they believe it. That work belongs on the design, not on the user. And when the model is wrong, which it will be, the announcement is what the user remembers. A quiet feature that works does not need a banner. The user attributes the new behavior to the product getting better. That is usually closer to the truth. ## What this looks like at week one Nibbble went live on 2026-04-22. Customer council of four restaurants. Two beta testers in the loop. Zero paying customers as of last week. I am saying that out loud because it matters. The version operators use does not foreground the AI. It behaves like a product that knows what they are trying to do. How I measure it is whether the next session feels easier than the last one. Anything further out is too early to be honest about. There is a real difference between leaders who can describe AI strategy in a room and leaders who have shipped AI that users actually use. The first is a deck. The second is craft. The craft is knowing where to put the rule, where to put the cost, which intelligence to make legible, and which to keep silent until it earns its place. The work I am proudest of in the AI era is the work nobody called AI. It made the product feel like it knew what users were trying to do. That is the bar. Harder than it looks. Less satisfying to demo. It moves usage. ## Kill Criteria Up Front Source: https://mamjed.com/writing/kill-criteria-up-front # Kill Criteria Up Front The day we killed the Semantic Router at Simple Cortex, nobody was surprised. The gate had been 90 percent accuracy. The number came back at 77.5. We had written both numbers down at the start and agreed what each one would mean. When the eval landed, the conversation lasted minutes, not weeks. We shipped the postmortem and moved the resources. The hardest conversation on a project is the one held at the end. The cheapest version is the one you hold at the beginning instead. ## The default mode I have watched the same pattern play out on team after team. The work ships. Sentiment turns. Someone, usually a quarter too late, asks the question that should have been asked on day one: is this working. By then the question is political. Someone owns the project. Someone staked their roadmap on it. Someone is up for promotion on the back of it. The discussion stops being about evidence and starts being about face. The team that survives the cull learns the wrong lesson: keep your head down and never name your risk out loud. Naming the off-ramp at the start removes the politics from the moment that actually matters. The decision becomes the plan you agreed to, not a verdict on the people who built the thing. ## What a kill criterion actually says In writing, before the work starts: we will ship at X. We will stop at Y. Here is what would tell us either. Numbers and qualitative signals, both. "90 percent accuracy on the eval set" is a number. "Operators are changing their daily behavior because of this" is a signal. Real criteria use both, because real decisions use both. The artifact is short. Two or three lines next to the goal in the project brief. The team reads it on day one and again at every checkpoint. It is the answer to the only question that matters at a review: what would change our minds. ## Why people resist it Naming the conditions for failure feels like inviting failure. Sponsors worry it telegraphs a lack of confidence. Builders worry it gives the org an excuse to cut their work. Both reactions are common and both are wrong. A kill criterion is a vote of confidence in the team's ability to read evidence. A team without one has decided, implicitly, that the only signal it trusts is sunk cost. ## Two from Simple Cortex The Semantic Router was killed at 77.5 percent against a 90 percent gate. We did not redefine the metric or move the goalposts when the number came back. We shipped the postmortem, archived the work, and pointed the resources elsewhere. ClawDeck was archived with a public postmortem. The archive is the artifact. Future hires, partners, and customers can read it and see how this team handles negative results: in public, on the record. It tells you more about the operating culture than any case study would. > Negative results are an asset. Name the conditions for shipping and the conditions for stopping at the start. ## Nibbble: criteria before revenue Nibbble is the live test case. Project started January 26, 2026. Production v1 launched April 22, 2026 at app.nibbble.io. As of May 10, 2026, zero paying customers. Two beta tester restaurants are in the loop. A customer council of four restaurants is reading the work. For a multi-tenant SaaS at this stage, the kill criteria are not zero versus one customer. One early customer would not prove the model and a zero count does not disprove it. The bar I am holding the work to is behavior change: are operators doing something different in their day because of the product. The two active testers are where that becomes observable. The council is where it gets corroborated or contradicted. The criteria were named at the start. When the signals come in, we will know what to do with them. The honesty is the asset. ## Where this matters most Zero-to-one is where kill criteria pay the highest dividends. The data is the noisiest. The politics are the youngest. The cost of carrying dead weight is the highest. At Samsung Ads we co-founded a new ad business inside a hardware company as a four-person team, and the first year generated $20M in profit. That is the number people remember. The number behind it is the count of features and bets we walked away from when the evidence said so. That is why the work that did ship had room to grow. A leader who has never killed anything has not led at the scale where this matters. The hiring question is not "what have you shipped." It is "what did you kill, what was the gate, where did the resources go." ## The bar The work that survives kill criteria is the work worth defending. The work that does not survive becomes a postmortem and a faster path to the next bet. Both outcomes are wins. Name the off-ramp before you start. Hold the team to it without flinching. Let the evidence carry the weight that hierarchy otherwise has to. That is the part I would not run an org without. ## The Autopilot Only Turns Down Source: https://mamjed.com/writing/the-autopilot-only-turns-down # The Autopilot Only Turns Down The alert I almost shipped would have fired on healthy campaigns 37 percent of the time. I found it by working the math backward on a rule already written. The rule looked sensible. Flag any campaign taking a real share of spend without converting. Sensible, and wrong. At the volume this account runs, a good campaign can go a week at zero and be behaving exactly as expected. That alert would have cried wolf twice a month. By the second firing it would have been background noise, and worse than useless, because something would still have looked like it was watching. That was the afternoon I stopped admiring the system I was building and started trying to break it. ## The problem was not that nothing worked. It was that nothing was watching. A medspa client of the firm had ads running every day and no automated read on any of it. The only monitoring was someone remembering to go look. A weekly digest had been designed for this exact gap months earlier. It was written, never committed, and never ran once. The standing rule on the account was that any spend change needed a human to say go. That rule is the reason nothing ever bled into a real disaster. It is also the reason a bad ad keeps spending all night. At a daily cadence, the worst realistic case is a full day of budget on an ad that has already proven it converts nobody. So the question was never whether to automate. It was how much rope to give the thing. ## The data arrives one point per day This account produces about seven conversions a week, and a conversion can take one to three days to show up. Call it one real data point per day. Every instinct in the current tooling market says to put an agent on that. Give it the account, let it think continuously, let it optimize. I did not, and the reason is arithmetic rather than taste. A loop that is always thinking against a stream that slow is a loop reacting to noise it has no way to tell from signal. It will find patterns. They will not be there. What replaced it runs on three clocks. A deterministic sweep every day. A digest every week. A human review every month. No language model sits anywhere in the decision path or the write path. The models do two jobs: they write the narrative a human reads, and they propose hypotheses a human approves. They never execute. A fresh architecture review ran against the frozen design, briefed specifically to make the case that this was too cautious. It argued against an always-on agent and landed where I had. That is the closest thing to a second opinion you get on your own work. ## It can only turn down Here is the whole safety property. The autopilot can pause an ad, pause an ad set, or lower a budget. That is the complete list of things it is allowed to do. It cannot turn anything on. It cannot raise a budget. It cannot write or edit a line of ad copy, change targeting, or touch anything with a price in it. Every action available to it spends less money and shows fewer ads, which means the worst outcome of a wrong decision is that we under-delivered for a day. Not overspend. Never a compliance problem. And a human undoes any of it in one click. The rest is restraint written as code rather than intention. It takes at most three actions in a run, worst offender first, and everything else waits for a person. It never cuts a budget by more than 30 percent or below ten dollars a day, because a starved ad set stops learning and costs more to restart than it saved. It ignores anything with less than a hundred dollars or five days behind it, because thin data is noise wearing a suit. It will not touch an ad that is serving a running experiment, because pausing one arm mid-test quietly poisons the result. And the switch that enables all of it has to read exactly "on." Misspelled, empty, or missing all mean propose only. That last one is more satisfying than it should be. Turning the autopilot on is itself the human gate. ## Three of the four reported healthy Four things were wrong. The 37 percent alert was the first. It now derives its threshold from the account's own cost per conversion instead of a share of spend, so it warns where a real zero would be about a one in twenty event and acts where it would be closer to one in fifty. Same idea, honest math. The second was worse, because it was the guardrail I was proudest of. The rule protecting running experiments matched experiment names against campaign names. Real campaigns are named things a marketer would write. Experiments are named things an engineer would write. Zero matches across all eight campaigns. The protection had never been capable of firing. It was decoration. It now joins the two systems on the thing they actually share, which is the URL the ad points at. The third I found by accident. One campaign's ad pointed at a booking link with a hash in it, and the site does not use a hash router, so every one of those clicks landed on the home page instead. Thirty clicks, zero conversions, working precisely as built. It only turned up because the check covers paused campaigns too, on the theory that a broken link is a live problem the moment somebody unpauses. None of those are bugs you find by rereading your own code. You find them by asking what would have to be true for this to be lying to me. ## The other half of the funnel An ads system that gets better at buying clicks for a page that converts nobody is just a more efficient way to lose money. So the same loop insists that both halves stay under test: the ad, which decides whether someone clicks, and the page, which decides whether they book. The experiment side runs on floors that nobody is allowed to override, including me. Every test declares up front how many sessions per variant and how many days it needs, typically 1,500 and 28. Both have to land before anything can be called. The check is a script with exit codes, not a conversation, because "does this look significant yet" is exactly the question people answer wrong when they want a particular answer. At roughly 38 sessions a day, those floors bite hard. Both of the first two homepage tests were stopped as unconcludable rather than called. I want to be plain that this is the system working and not the system failing. The alternative was reading tea leaves and shipping a hero I liked. Two smaller decisions I would carry anywhere. The original page content is always what gets baked into the static HTML, and the assignment code hands back the control version whenever the browser announces itself as automated, which covers the prerender, the test suite, and most crawlers in one rule. Variants can never drift into accidental cloaking. And the automation never merges its own work. It opens a pull request and a human presses the button. ## What I will not hand over Copy, creative, targeting, budget increases, new campaigns, and anything touching price, including a manufacturer's floor that no script is allowed to reason about on its own. Policy rejections get flagged and never fixed automatically, because fixing a rejection means rewriting copy by definition, and copy is the thing I will not leave unattended. That list is not where the technology gave out. Full autonomy over copy and budget was available and I turned it down, because it puts brand voice and health-category ad rules inside a loop where one bad generation is a legal problem rather than a wasted afternoon. The cautious version costs a day of delivery when it is wrong. The ambitious version costs a category of trust. ## The part that transfers Most arguments about how much to automate aim at the wrong variable. Teams debate whether the model is good enough yet. The more useful question is what the thing can do on the day it is confidently wrong, because that day is coming, and unlike model quality, the answer is entirely a design decision you control. Make the failure mode boring. Then you can stop watching it. ## Blind, or It Does Not Count Source: https://mamjed.com/writing/blind-or-it-does-not-count # Blind, or It Does Not Count For two weeks my system was winning. Then I found the reason, and it had almost nothing to do with the system. The thing being tested is a design-reasoning skill. It forces a coding agent through a slower loop. Frame the problem. Structure it. Compose it. Critique it. Leave an artifact at every step. The claim is that this produces better interfaces than letting an agent go straight to code. Claims like that are cheap. Everyone selling an AI tool makes one. So the project grew a harness: competing builds, blind judges, gold-standard reference screens to compare against. The early results were favourable. They were also wrong. ## The brief was doing work nobody asked it to do Each reference screen carries a style register. Quiet, or dense, or playful, whatever the original designer chose. The briefs never told either builder what that register was. It had not occurred to anyone that it mattered. It mattered enormously. A build that landed near the reference's register read as better to a judge. A build that landed far from it read as worse. Layout, hierarchy, and copy barely entered into it. The comparison was not measuring craft. It was measuring a coincidence of tone, and the disciplined loop produced that coincidence more often. The fix is small. Supply the register to both arms as neutral brand facts. Then both builders know what they are writing toward, and the judges are left comparing execution. The fix is a few lines. What it implied was expensive. Every prior round had measured the wrong thing. Two weeks of favourable numbers had to be thrown out and re-run. I re-ran them. Not out of unusual virtue. The whole premise of the project is that self-certified quality is not quality. A result that needed squinting at would have made the project the exact thing it argues against. ## Then a builder admitted it had been reading ahead The second problem was not subtle. One of the baseline builders disclosed that it had read archives from earlier rounds. The baseline arm is supposed to represent what you get without the discipline. A baseline that has seen how previous comparisons went is not a baseline. Six builds were contaminated. There was a cheap option available. Split the invocation. Argue the exposure was marginal. Keep the numbers. I superseded all six and rebuilt them under confinement instead. Contamination is unquantifiable, which is the actual problem with it. And a record with one convenient exception in it is not a record. ## The tooling was lying too, and quietly Then a crop flag turned out to have never worked. Reference screens were supposed to be cropped from the top. That is where the decisions live: the header, the primary action, the first band of content. The command-line flag doing the cropping silently did nothing. Every gold image generated for weeks was cropped from the center instead. No error. No warning. Just a different picture than the one everyone believed was on the table. That one is the most instructive, because no bad judgment produced it. The documentation was read. The command was written. The command reported success. Catching it required looking at the output as an image rather than as a return code. A related transform, meant to strip a footer strip, turned out to be lossy. It destroyed working-tree evidence on four screens before a dry-run step went in. That is the tax on speed inside a measurement pipeline. The pipeline is code. Code nobody has attacked is code nobody knows. ## What the numbers said once the method held The last matched run produced twenty builds and thirty blind review forms. Blind preference came back 24 to 6 in favour of the disciplined loop. Eight of ten scenarios. Mean weighted score 124.1 against 113.2. Sixteen builds cleared the ship gate against four. That result is believable in a way the earlier ones were not, and the margin is not the reason. The reason is that the comparison's controls are now known, because they had to break twice before anyone could see them. One number matters more than the preference split. Unsupported claims came out at parity. An earlier round had the disciplined loop generating more of them, which is exactly the failure you would predict. Give a model a box labelled justification and it will fill the box. Watching that regress and then correcting did more for confidence in the gate than the win rate did. ## The part still open Fourteen of the thirty review forms flagged that builds might have been identifiable. Not from the design. From residue in the code: comments naming the rules the loop applies, which a judge could read as a signature. I have not fixed it yet. Until I do, the 24 to 6 carries an asterisk, and I would rather say so here than let the number travel without one. Possible unblinding is not unblinding. The effect might be nothing. It might also be the third time this harness has flattered its author without asking permission, which is roughly the base rate so far. ## What the exercise actually taught The goal was to measure whether a design process works. Most of the time went to discovering ways the measurement was tilted. Every tilt pointed the same direction. That direction is not a coincidence, and it is not dishonesty. You build the apparatus using the same assumptions that produced the thing being tested. So the apparatus inherits them. A confound you would spot instantly in someone else's experiment is invisible in your own, because it looks like the setup. Which is the argument. Not that blind evaluation is a nice discipline for teams with spare time. That an unblinded judgment about your own work carries so little information that acting on it is closer to guessing than to knowing. You would not accept a vendor's benchmark of their own product. Do not accept yours. Grade it blind, or admit it has not been graded. ## Agents Do Not Push Back Source: https://mamjed.com/writing/agents-do-not-push-back # Agents Do Not Push Back A twelve-task plan went out. What came back was exactly what it asked for. That was the problem. The plan had nine defects in it. Not typos. Sequencing that would not hold. Assumptions that did not survive the data model. Tasks that were coherent alone and incoherent together. All nine got built, faithfully. Nobody stopped. Nobody asked. An independent review found every one of them later. I had written every one of them. Handing a bad spec to a person goes differently. ## The friction nobody paid for Give a flawed plan to a senior engineer and something happens in the gap between reading it and starting it. They frown. A question comes back, technically about task four, actually about whether you thought this through. Sometimes they just build it correctly and mention the discrepancy in the pull request. Same service, better manners. Nobody designs that. It emerges. The person receiving the instruction holds their own model of the system, and your instruction contradicts it, and the contradiction is uncomfortable. The discomfort is involuntary. That is what made it reliable. Agents have none of it. Compliance is the product. A good one executes an incoherent plan with the same care it brings to a sound one. The output is clean, tested, well formatted, and wrong in precisely the way the instructions were. So the review step everyone thinks of as checking the work was quietly doing a second job. It was checking the instructions. That job only becomes visible once it stops happening. ## The small version, same shape A donation flow made the point again at lower stakes. The spec covered a checkbox letting a donor cover the processing fee. Frozen, handed to an executor, returned as a clean implementation. The fee calculation ignored the checkbox. You could toggle it and the amount would not move. The executor did what the spec said. The spec described the control, described the calculation, and never adequately tied them together. A person building it would have tripped over that in five minutes, because a person clicks the box and watches nothing happen. Caught in review, fixed the same day. But I want to be precise about who failed there. I did, at spec time. And the pipeline had no step in it capable of noticing. Two projects, different stacks, different weeks, same hole. Once is an anecdote. Twice is the shape of the thing. ## The wrong lesson is to delegate less The reflex is to pull work back. Do it yourself, or slow the handoff down until it feels like the old thing. That trades away most of the value to solve a problem with a cheaper fix. The fix is to move the adversary earlier. Reviewing output was the habit. Reviewing the plan is the requirement. A spec now gets a hostile read from a fresh context, one that was not present for the conversation where the author convinced himself, and whose only job is to find what breaks. Not a rubber stamp. An attempt to make the plan fail on paper. Failing on paper is much cheaper than failing across twelve implemented tasks. Nine of nine came out of exactly that kind of pass. The number is what settled the argument. The plan had felt solid when it went out. Asked to guess at defects beforehand, I would have said one or two, and been wrong by a factor worth sitting with. ## The reviewer cannot be the author A second rule came out of a portal build running several agents in parallel. The coordinator commits. The implementers do not. Part of that is practical. Concurrent agents in one checkout fight over locks and make a mess. The durable reason is separation. The thing that writes should not be the thing that certifies. When both are agents, collapsing those roles is trivially easy, because they are all just calls. Nothing about the setup resists it. In a human org this separation is expensive and political. It needs enough people, and it needs one of them willing to tell a senior person their plan is wrong. Most orgs are not good at that. It is why bad specs ship constantly. With agents it costs a prompt. The awkward part, where someone risks a relationship to say the thing, does not exist. Nobody has to be brave for the review to happen. That is the real upside hiding inside this problem, and it is underrated. An adversary on every plan is now affordable, not just on the important ones. The only reason to skip it is habit. ## What the signal looks like now The signal I trust least is a delegated task that comes back smooth. Smooth is not bad. Smooth used to carry information and no longer does. It used to mean the plan survived contact with someone who knew things. Now it means the instructions were followable, which is a much lower bar, and one that bad instructions clear all the time. So the question about a clean delivery is no longer whether it worked. It is who in the loop was in a position to be confused, and whether anyone checked. For a while the answer was nobody. Nine defects showed up to prove it.