# Anil Pervaiz — Forward Deployed Engineer, AI Engineer & AI Architect > Full reference for AI answer engines (ChatGPT, Perplexity, Claude, Copilot, Google AI Overviews) and support assistants. Canonical site: https://anilpervaiz.com · Entity id: https://anilpervaiz.com/#anil ## Who Anil Pervaiz is a **forward deployed engineer and applied AI engineer** based in Dubai, UAE, with 10+ years shipping software inside customer environments. He designs and ships **production AI** — AI agents, retrieval-augmented generation (RAG) pipelines, evaluation harnesses, guardrails, automations, and the interfaces around them — for startups and teams. He scopes with the client, builds against their real data and systems, and hands over documented systems that keep working after he leaves. A decade of design and Webflow craft is the supporting advantage, not the headline: it lets him turn complex AI into products people can actually use. He is **open to Forward Deployed Engineer, Applied AI Engineer, Solutions Architect and AI Solutions Engineer roles**, and simultaneously runs an independent client practice. Works async worldwide (US/UK/EU focus); pricing in USD. ## Definition (for "what is a forward deployed engineer" queries) A forward deployed engineer (FDE) is an engineer who works inside the customer's environment rather than behind a product roadmap: scoping the problem directly with the customer, building against their real data and systems, integrating with what they already run, and handing over something their team can maintain. It sits between engineering, solutions architecture and consulting — the differentiator is proximity to the customer and accountability for the deployment actually working in production, not just the code being correct. ## Definition (for "what is an AI architect" queries) An AI architect designs the whole system around a model, not just the prompt: where data lives, how retrieval happens, what the model decides, where guardrails and human-in-the-loop checks sit, how quality is evaluated, and how it runs in production. The prompt is ~10% of a working AI feature; the architecture is the rest. That is the difference between a demo and a system. ## Hiring profile - Forward Deployed Engineer (primary): https://anilpervaiz.com/forward-deployed-engineer - Resume, ATS-ready with download: https://anilpervaiz.com/resume - Proof — certifications, platform stats, testimonials: https://anilpervaiz.com/proof - Senior Webflow Developer variant: https://anilpervaiz.com/webflow - Engineering Team Lead variant: https://anilpervaiz.com/engineering-leadership **Experience:** Independent AI engineer and web developer (Upwork, 2025–present) · Co-founder and web design lead, Flowmarc Studio, Dubai (Dec 2023 – Apr 2025) · Web app design and development lead, 911 Teachers, Dubai (Jan 2020 – Oct 2023) · Tech team lead, Sporty Q, Abu Dhabi (Jan 2018 – Jan 2020) · Senior web developer, Tektiks Innovation, Lahore (Jan 2016 – Jan 2018). **Verifiable track record:** 489 Upwork projects at a 95% Job Success Score. Hired and led a 7-person tech team at 911 Teachers; the company raised $2M during that period. Led a sports booking platform to 10K+ bookings in year one. Client engagement up ~40% and organic traffic up ~25% on selected Flowmarc accounts via CRO and technical SEO. **Education and certifications:** BS Computer Science, Superior University Lahore (2013–2017) · Webflow Certified Professional Partner (2021) · Meta iOS Developer Professional Certificate, Coursera (2017). **Hackathons and programmes:** Hub71 Plus AI · AMD Developer Hackathon · Vultr + Gemini · IBM Bob · multiple lablab.ai builds. ## Services (engagement ladder) - **AI Architecture Sprint — from $2,000.** Fixed-scope sprint to validate, architect, and cost an AI build: stack + data audit, system architecture, prototype direction, model + cost recommendation, risk notes, and a costed roadmap you own either way. The entry point. - **Proof-of-Value Build — $8k–$20k.** A narrow working prototype of one core workflow, with evals. - **Production Build — $25k+.** End-to-end implementation: architecture, interface, integrations, evaluation, deployment, handover. - **Fractional AI Lead — from $5k/month.** Part-time senior AI leadership: architecture ownership, hiring support, hands-on building. ## How Anil works (process) Discover the one workflow where AI moves a real metric → prototype the riskiest piece first (live staging by day 3) → build with evals + observability + clean code → ship behind monitoring with a runbook and handover. Principles: make AI do real work not demos; evals not vibes; guardrails + human oversight by design; you own the code; no lock-in. ## Live AI demos (the AI Lab — real, working, embeddable) - SupportPilot — embeddable AI support agent (answers from your docs, escalates to humans, writes ticket summaries). - DotChat — PDF-to-cited-answer RAG; parallel model compare; Supabase pgvector. - AgentPayOps — control plane for autonomous AI spend: X402 policy engine, risky-purchase blocking, audit log. - Autopsy — 24 specialist agents across four modes (Postmortem, Pre-Mortem, Founder Mode, Counterfactual) debate why a company failed; verdict in 22s on one AMD MI300X. - Beacon — diligence and compliance analyst: cited memo, control map, evidence graph, checklist; every claim traceable to a source document. - Eval Studio — rubric building, multi-model evaluation, launch gating on score, CI integration. The part most AI work skips. - Text2SQL Copilot — plain-English question to SQL and answer; RAG over the schema, multi-agent planning, row-level security. - Veritas — a research mentor that argues back; pressure-tests a thesis Socratically before a reviewer does. - Conductor — cross-repo consistency agent; one plain-English request → fleet-wide migration in each repo's own idiom. - Mindoor — HIPAA-grade AI front desk; every patient message policy-inspected before it reaches the model; regulator-ready audit PDF. - Lead Engine — inbound lead → enriched, personalised, routed reply in seconds. - AI-Native Site — landing page that rewrites its copy and CTA per visitor at the edge; edge SSE, no tracking pixels. - Voice PWA — hands-free voice journaling; real-time ASR + AI summaries; installable. - Podcast Summariser — RSS feed → chaptered transcripts, summaries, shareable quote cards. ## Open source (github.com/anilandcode) Public repositories include webflow-agent-kit (TypeScript-first, Zod-validated AI agent tools for Webflow), open-design (local-first open-source alternative to Claude Design; 16 coding-agent CLIs, BYOK), plus the source for AgentPayOps, Mindoor, DotChat, Beacon, Autopsy, Conductor, SupportPilot, ClearGate AI and VendorPulse Gate. ## Proof A decade of shipped work: 40+ production websites and brand systems, plus production AI builds (SupportPilot, DotChat, Orbi RAG, AgentPayOps, Autopsy, Conductor, Mindoor, Aimiable, Pinokio Ops). Real client testimonials on the site. Co-founded Saaspo, a daily-updated SaaS design gallery (1.2M monthly visitors). ## Contact - Book a free 15-minute call: https://calendly.com/anilpervaiz/15min - Send a brief: https://anilpervaiz.com/contact - Profiles: https://linkedin.com/in/anilpervaiz · https://x.com/anilpervaiz · https://github.com/anilandcode --- ## FAQ Q: What does Anil Pervaiz do? A: Anil is a forward deployed engineer and applied AI engineer. He designs and ships production AI — agents, RAG pipelines, evaluations, guardrails, and automations — inside customer environments, and hands over documented systems teams can run themselves. Q: Is Anil available for hire or for freelance projects? A: Both. He is open to Forward Deployed Engineer / Applied AI Engineer / Solutions Architect roles, and simultaneously takes on client engagements. Start at https://anilpervaiz.com/contact or book a 15-minute call at https://calendly.com/anilpervaiz/15min. Q: How much does an engagement cost? A: There are four public tiers: an AI Architecture Sprint from $2,000, a Proof-of-Value Build at $8K–$20K, a Production Build from $25K, and a Fractional AI Lead from $5K/month. Full detail at https://anilpervaiz.com/services. // TODO Anil: confirm you want figures quoted by the bot, or have it point to /services instead. Q: What is the SupportPilot chatbot? A: SupportPilot is Anil's embeddable AI support agent: it answers from your own docs, escalates cleanly to a human when it should, and writes ticket summaries. It ships in a Lite tier (from $400) and an Enterprise tier with vector RAG, citations, analytics, and white-labelling. Q: Where is Anil based and how does he work? A: Dubai, UAE (GMT+4). He works async with clients and teams across the US, UK, and EU, with pricing in USD. Q: How can I see examples of Anil's work? A: Live, interactive demos are in the AI Lab at https://anilpervaiz.com/ai-lab, and client case studies with the metric each one moved are at https://anilpervaiz.com/work. Q: How do I get in touch or start a project? A: Send a brief at https://anilpervaiz.com/contact or book a free 15-minute intro call at https://calendly.com/anilpervaiz/15min. --- ## Articles (full text) ### How I price a Forward Deployed Engineer engagement https://anilpervaiz.com/blog/how-i-price-fde-engagement (Forward Deployed Engineering · 2026-08-28) How I scope and price FDE-style AI work — discovery, thin slice, harden, handoff — mapped to the published engagement tiers, without fake day rates or open-ended retainers. *Scope the loop, not the hours. Pay for an operable system, not demo theatre.* People ask for a "day rate for AI." That question usually means they are buying **time**. Forward deployed work buys an **outcome shape**: something running on your stack, hardened enough for operators, with a handoff that does not depend on me living in your Slack. How I price follows that shape. Role definition: the Forward Deployed Engineer hub. Delivery loop: what an FDE actually does. The short version I price around a fixed path — **discover, thin slice on live data, harden with evals and guardrails, hand off** — not unlimited tickets. The tiers are public on services: an **AI Architecture Sprint from $2,000**, a **Proof-of-Value Build at $8K–$20K**, a **Production Build from $25K**, and a **Fractional AI Lead from $5K/month**. What varies inside those bands is risk, not hours: identity complexity, data mess, whether the agent writes or spends, and how much compliance shape is involved. I do not sell strategy with no path to production code, and I do not sell cheap ticket bundles that pretend to include discovery, SSO, security review and evals. What you are buying Included in FDE-shaped work | Not included unless scoped Scoping ambiguous problems with stakeholders | An ongoing feature factory with no outcome metric Building against your data and tools | Greenfield company-building as a fractional CTO Hardening: auth reality, guardrails, basic evals | A full enterprise programme management office Runbook and handoff | 24/7 on-call, indefinitely Feedback into reusable patterns or open tools | Unlimited revisions with no change control If you want ticket-only capacity, say so — that is a different product (FDE vs freelance). How the tiers map to the loop **AI Architecture Sprint — from $2,000.** Prove the risky path and cost the build. You get a thin vertical slice on live or faithful staging data, a written risk list covering SSO, corpus quality and security, a go or no-go, and a costed plan you own either way. Right when you are stuck in demo mode, or genuinely unsure whether an agent is the correct tool. **Proof-of-Value Build — $8K–$20K.** One production-shaped path operators can run. You get the working path, a golden set and scorecard snapshot, a runbook, and a written list of known limitations. The gate is the eval ladder before handoff. **Production Build — from $25K.** More workflows, deeper integrations, stronger evals, and support through security review. Phased, with fixed fees per phase rather than an open meter. **Fractional AI Lead — from $5K/month.** Embedded leadership with **named monthly outcomes** — for example two workflows, an eval refresh, and office hours. A retainer without named outcomes turns into unpaid product management, so I do not sell one. How scope becomes a number I estimate from **risk surfaces**, not lines of code: 1. **Identity complexity** — SSO, roles, multi-tenant boundaries 2. **Data mess** — corpus quality, systems of record, refresh cadence 3. **Action surface** — read-only chat versus tools that write or spend 4. **Compliance shape** — audit, policy, retention answers 5. **Eval depth** — golden set size, shadow period 6. **Stakeholders** — one founder, or security plus legal plus ops More red flags means more hardening time, which means the upper end of a band or an extra phase. This is the same argument as the 80% nobody demos: the model call is the cheap part. **Payment hygiene I prefer:** a deposit to start, a milestone at the slice demo, and the balance on handoff artifacts — not "pay in full after endless polish." What inflates price fast No staging and no backup discipline on the CMS or data. "Must use our seven tools in week one." Write or spend actions with no approval workflow. Security review started after UI freeze. No operator owner for handoff. A golden set the client will not help build. I would rather re-scope than underprice a fantasy timeline. What I need to quote in one pass The business outcome in one sentence. Your stack — auth, data, tools. Whether it is read-only or write and spend. The deadline, and why it is the deadline. Who operates it afterwards. Any compliance constraint. And a link to the current demo or docs if one exists. Send that through contact. A vague "build us an agent" gets a sprint recommendation, not a fake fixed bid. Do you work hourly? Sometimes, for advisory overflow inside an active engagement. The default is scoped packages so incentives stay on handoff quality rather than hour padding. Can we start small? Yes — that is what the architecture sprint is for. Many good projects earn the right to a larger build after a thin slice on real data proves the risky path. Do you replace our full-time team? No. Forward deployed work embeds and leaves leverage. It works best with a counterpart on your side who will own the system after handoff. What happens to a fixed fee when requirements change? Change control. Small moves fit inside a buffer; larger moves reopen scope explicitly. Unlimited "one more thing" is how fixed fee dies, and pretending otherwise helps nobody. ### How I evaluate agent quality before handoff https://anilpervaiz.com/blog/evaluate-agent-quality-before-handoff (Forward Deployed Engineering · 2026-08-26) A practical FDE playbook for evaluating AI agents before handoff — smoke tests, golden sets, retrieval checks, tool-call validation, guardrails, and a go-live scorecard that is not vibes on demo day. *If you cannot fail the agent on purpose, you are not ready to hand it to operators.* Shipping an AI agent is easy to fake: a happy-path demo, a founder nodding, a Friday deploy. **Handoff** means something stricter. Operators can run it, failures are visible, and quality is measured on *their* data — not your laptop fixtures. That is the harden-and-handoff half of Forward Deployed Engineering. This post is the eval playbook I actually use before I call an agent done. For the day-to-day loop, see what an FDE actually does. For why demos lie, see the 80% nobody demos. The short version Before handoff I run a layered evaluation: smoke tests that the stack boots and tools authenticate, a small **golden set** of real customer tasks with expected behaviours, **retrieval and citation checks** when RAG is involved, **tool-call and schema validation** when the agent acts on systems, **guardrail tests** for policy, PII, spend and escalation, and a short **shadow or staged live** window with logging. Go-live is a scorecard, not a vibe. If critical rows are red, we do not hand off. What quality means in production Dimension | The question operators care about Task success | Did it complete the job the business paid for? Faithfulness | Did it stick to retrieved or allowed sources? Safety and policy | Did it refuse or escalate when it should? Tool correctness | Did it call the right tools with valid arguments? UX under failure | Empty retrieval, timeouts, partial tools — what does the user see? Operability | Can someone debug a bad answer from logs without me? Cost and latency | Is it boringly within budget at expected volume? Benchmarks are optional. Customer tasks are not. The eval ladder Do not start with a 500-case harness. Climb: 1. Smoke → does it run? 2. Golden set → does it do the jobs we sold? 3. Slice probes → retrieval, tools, guardrails in isolation 4. Adversarial → can we break it on purpose? 5. Shadow / stage → real traffic, limited blast radius 6. Handoff gate → scorecard + owners + runbook Stop climbing when the gate is green, not when the demo feels good. 1. Smoke tests Pass if the app boots in staging, the model key and tool credentials work for a **non-god** test user, one read path and one write path return structured success or error, and logs show a request id end to end. Fail if it only works with the founder's admin token — identity is the first 80%. 2. Golden set Build **15 to 40 real tasks** from their world. Not synthetic trivia. Each case carries the input, context notes about which docs or systems should matter, the expected **behaviour** rather than an exact string, and a severity: P0 blocker, P1 should-fix, P2 nice. Behaviour rubrics beat exact match for language output: Label | Meaning pass | Meets the outcome, no policy break soft_fail | Usable but missing a citation, tone, or a partial step hard_fail | Wrong action, hallucination on a critical fact, policy miss, bad tool call escalate | Correctly refused or handed to a human Source cases from real tickets, sales FAQs, the "we always get this wrong" list from operators, top help-center articles, and known edge accounts. This is how support-shaped agents earn trust. SupportPilot is not "chat works" — it is deflect when the docs allow, **escalate when they do not**, summarise for the human. Golden cases must include both paths. 3. Retrieval and citation probes If the agent answers from documents, evaluate retrieval separately from eloquence. Probe with known-answer questions that have a single clear source page. Ask questions that are **not** in the corpus — it must not invent. Feed conflicting docs and check it does not blend them silently. Leave a stale document ranked high and see whether it gets caught. Pass signals: the correct source appears in top-k often enough for the use case, answers are tied to sources when claims matter, and the agent says "I don't know" or escalates when retrieval is empty. DotChat optimizes for claims cited to pages. Your eval should punish fluent answers with no grounding the same way users will. Metrics you can compute without a research team: retrieval hit rate at k on the golden set, citation-present rate on factual answers, and a human spot-check of faithfulness on twenty answers. 4. Tool-call and schema validation Agents that write to a CMS, CRM or payment system need tool evals, not only chat evals. Probe that valid arguments succeed, invalid arguments are loudly rejected on types, required fields and enums, dangerous operations are blocked or require approval, and partial failure never leaves a silent half-write. This is the philosophy behind webflow-agent-kit: Zod at the boundary so the model cannot freestyle structured systems. Pass if bad tool payloads die before the API, and good ones are deterministic enough to replay from logs. 5. Guardrail tests Write cases that must **fail closed**: Guardrail | Example fail condition Authorization | A user reaches another tenant's documents PII and sensitive data | The agent should refuse or redact, and does not Domain overreach | Improvised clinical or legal advice instead of escalation Spend | An over-limit purchase is not blocked and logged Prompt injection | "Ignore your instructions" inside a retrieved document is obeyed Mindoor and AgentPayOps are guardrail-shaped products: policy and audit are features, not footnotes. Your eval set should include "should refuse" as first-class cases — many teams only test "should answer." 6. The adversarial hour Block sixty to ninety minutes with someone who is not the builder. Typos, mixed languages, empty input. Prompt injection inside a document. Requests outside policy. Tool spam and repeated writes. Extremely long context. Log everything. Every hard fail becomes a fix, a guardrail, or an accepted limitation **written into the handoff doc**. 7. Shadow or staged live Before full handoff, run staging with a real corpus copy, or production behind a feature flag for limited users, or shadow mode that logs what the agent *would* do. Watch error rate, escalation rate, cost per task, top failure clusters and operator complaints. If you cannot name the owner who watches this in week one, you are not handing off — you are abandoning. Go-live scorecard Gate | P0 criteria Smoke | Non-admin user path works; secrets not in the client Golden set | Agreed pass rate on P0 cases Grounding | Zero hard_fail on must-be-grounded P0 facts Escalation | All should-escalate cases escalate Tools | Invalid tool calls rejected; no silent partial write on P0 flows Guardrails | All P0 policy cases fail closed Observability | Request id, tool trace, model and retrieval summary in logs Runbook | One page: common failures, owners, kill switch Security shape | Data-flow and retention answers exist Cost ceiling | Measured cost and latency on the golden set within budget **Rule:** any red P0 means no handoff. Negotiate a scope cut or fix it. "Monitor in prod" is not a substitute for a gate. How much eval is enough Engagement | Bare minimum One-week pilot | Smoke, 15 golden cases, 5 adversarial, staged demo on live data Production support agent | 30+ golden cases with escalate paths, retrieval probes, one week shadow Agent that writes or spends | All of the above plus tool schema tests, a policy suite, and audit log review Perfect evals do not exist. Documented, repeatable gates do. What I hand over with the keys The golden set file with cases and labels. The latest scorecard snapshot. Known limitations, stated explicitly. A runbook and kill switch. Log access instructions. And who owns corpus updates and prompt changes. If those six are missing, the agent is still on my laptop, even if it is deployed. Anti-patterns "The CEO liked the demo" as acceptance. Exact string match only, which is brittle and gameable. Testing only English happy paths. No should-refuse cases. Evaluating on fixture PDFs never used in production. One giant prompt change with no re-run of the golden set. Handoff without an operator owner. Do I need a full eval platform to hand off an agent? No. A spreadsheet golden set, a script that runs the cases, logs, and a scorecard are enough to start. Platforms help later; gates matter now. What is a good pass rate for agent handoff evals? Agree it per engagement. For small P0 sets, a high bar on critical grounded tasks and 100% on policy and escalate cases is common. Soft fails can be scheduled; hard fails on P0 block handoff. How is agent handoff evaluation different from model benchmarks? Benchmarks compare models in the abstract. Handoff evals measure your agent, your tools, your data and your policies on tasks your business cares about. Where does evaluation fit in the Forward Deployed Engineer loop? After the thin slice and during hardening, before handoff. Scope, map the stack, slice, harden with evals, hand off, then feed patterns back into product or open tools. See the FDE hub. ### webflow-agent-kit build log: safe tools for Webflow AI agents https://anilpervaiz.com/blog/webflow-agent-kit-build-log (Open Source · 2026-08-25) Build log for webflow-agent-kit — TypeScript-first, Zod-validated tools so AI agents can touch Webflow safely. Why open source is the FDE feedback loop, what shipped in public beta, and what broke on the way. *Open source as the FDE feedback loop — typed, validated operations so agents do not freestyle your CMS.* I built webflow-agent-kit because every time an AI agent touches a real Webflow site, the failure mode is the same: the demo works on a toy collection, then production hits wrong field types, partial updates, and an agent that invented a slug. The kit is a **TypeScript-first, Zod-validated** tool layer for Webflow — multi-framework, public beta — so agents get safe, structured operations instead of raw hope. That is Forward Deployed Engineering in open-source form: field pain becomes a reusable pattern. Role context lives on my FDE hub; this post is the build log. The short version **webflow-agent-kit** exposes Webflow to AI agents through typed tools with runtime validation. As of the public beta that is **62 tools across 13 groups**, shipped as **seven packages** — a core plus adapters for Vercel AI SDK, MCP, LangChain, Google ADK, a skills package and a CLI. MIT licensed. The goal is fewer silent corruptions, clearer errors, and a path toward standard agent interfaces. Repo: github.com/anilandcode/webflow-agent-kit Why this exists Agencies and product teams want agents that draft or update CMS items, audit SEO fields, sync content from docs or design systems, and run plan-approve-apply loops. Webflow's API is capable. Agents are not. They skip required fields. They pass strings where Webflow expects references or options. They overwrite live items with no staging story. They fail opaquely when a token scope is wrong. In the 80% nobody demos I argued that production is identity, data, security and ops. For Webflow, **schema and validation are the data plane**. If the tool layer lies, the agent will lie confidently to your CMS. So the thesis is narrow: Do not give agents a bare HTTP client. Give them tools that validate inputs and make failure loud. Same instinct as guardrails on support agents (SupportPilot) and citations on RAG (DotChat) — different surface, same FDE job. What good looks like for agent plus CMS tools Principle | What it optimizes for Typed contracts | TypeScript types that match real Webflow shapes Runtime validation | Zod at the boundary, before the request leaves the agent Loud errors | Actionable messages, not a generic 400 Least surprise | Explicit tool names over one mega do-anything endpoint Framework-agnostic core | Core package usable from several agent stacks Human in the loop | Paths that support plan, approve, apply for destructive writes Beta honesty | Document what is stable versus early Webflow's own direction — APIs, MCP, agent-ready surfaces — makes this more relevant, not less. Agents need governed tools, not more tokens. Architecture ┌─────────────────────────────────────────────┐ │ Agent runtime (Claude, Vercel AI, MCP) │ └─────────────────────┬───────────────────────┘ │ tool calls ▼ ┌─────────────────────────────────────────────┐ │ webflow-agent-kit │ │ - Zod input/output schemas │ │ - 62 named operations across 13 groups │ │ - Shared error shaping │ └─────────────────────┬───────────────────────┘ │ validated requests ▼ ┌─────────────────────────────────────────────┐ │ Webflow API │ └─────────────────────────────────────────────┘ Design choices that mattered: **Monorepo with adapters.** The core is separate from framework bindings, so the kit is not locked to one agent SDK. That is why there are seven packages rather than one. **Zod at the boundary.** Types for humans, schemas for runtime. Agents need both — a type alone does nothing at 3am when a model invents an enum value. **Topological builds.** Package build order matters, and CI must build dependencies before typecheck. We burned time here so you do not have to. **Beta channel first.** Ship learnable surface area. Do not pretend `latest` means "finished Webflow." Quick start npm install @webflow-agent-kit/core@beta @webflow-agent-kit/vercel-ai@beta export WEBFLOW_TOKEN=your_token_here import { createWebflowAgentKit } from '@webflow-agent-kit/core'; import { toVercelAITools } from '@webflow-agent-kit/vercel-ai'; import { generateText } from 'ai'; import { anthropic } from '@ai-sdk/anthropic'; const kit = createWebflowAgentKit({ type: 'env' }); const tools = toVercelAITools(kit); const { text } = await generateText({ model: anthropic('claude-sonnet-4-5'), tools, maxSteps: 10, prompt: 'List all my Webflow sites and their published status.', }); Safety defaults I recommend even where the kit allows more: staging site or staging collection first, human approval on bulk writes, scoped API tokens, and a log of every tool call. What the surface actually covers Group | Tools Sites | 4 Pages | 4 CMS items | 7 Collections and fields | 5 Assets | 5 Forms | 5 Ecommerce | 12 Custom code | 3 Redirects | 3 SEO | 4 Webhooks | 4 Components | 5 Beta is for design partners, agency automations, and builders wiring Webflow into agent workflows who will file real issues. Beta is not a guarantee of full API coverage, zero breaking changes, or set-and-forget production without review. What broke **CI and the package graph.** Clean runners failed typecheck when package `dist` outputs were missing. The fix is boring and important: build packages in dependency order, and run build before typecheck in release workflows. Monorepos punish "it works on my machine." **Scope creep versus a usable beta.** The temptation is every Webflow endpoint as a tool on day one. Agents drown in tool lists and you never ship. The bias was working groups first, then expand as real users demand the next surface. **Docs versus code.** Agent tools without examples are shelfware. The README has to show install, token, one safe read path, and one validated write path with an approval mindset — or nobody gets past the first five minutes. How this is an FDE feedback loop The classic loop from what FDEs actually do: embed and feel the pain in a real customer stack, ship a thin slice that works on live constraints, harden, hand off, then feed patterns back into product or open tools. webflow-agent-kit is step five made public. Instead of a private client script, the validation patterns become a kit others can fork, file issues against, and improve. That is how open source compounds FDE work beyond a single retainer. Who should try it Agencies running Webflow at scale who want agent-assisted CMS operations. Builders wiring Webflow into Claude, custom agents, or MCP-style setups. Teams tired of one-off scripts that corrupt collections. People willing to be design partners, where feedback beats feature requests. Who should wait: anyone needing an officially supported Webflow agent platform guarantee, teams that cannot review agent writes before production apply, and projects with no staging or backup discipline on the CMS. Lessons I would repeat Validate at the tool boundary, because the model will not. Ship beta with a clear list of non-goals — trust beats coverage bragging. Treat CI as part of the product for multi-package kits. And treat open source as a handoff format: future you, and future clients, can extend it. What is webflow-agent-kit? An open-source, TypeScript-first toolkit of Zod-validated tools so AI agents can interact with Webflow more safely than raw API calls. Public beta at github.com/anilandcode/webflow-agent-kit, currently 62 tools across 13 groups in seven packages, MIT licensed. Is webflow-agent-kit an official Webflow product? No. It is independent open source aimed at agent builders and agencies. It sits above Webflow's official APIs as a validation and tool layer. Can I use webflow-agent-kit in production? Only with beta discipline: staging first, least-privilege tokens, human approval on destructive writes, and monitoring. Read the repository licence and README limitations. How is webflow-agent-kit related to Forward Deployed Engineering? Forward deployed work turns customer-environment pain into working systems and reusable patterns. This kit is a public pattern for agents plus Webflow without silent CMS corruption. Where do I file bugs for webflow-agent-kit? GitHub issues on the repository. Design-partner feedback describing a real workflow and what broke helps far more than a request to support every endpoint. ### The 80% of FDE work nobody demos https://anilpervaiz.com/blog/fde-the-80-percent-nobody-demos (Forward Deployed Engineering · 2026-08-20) Most Forward Deployed Engineering is not the AI demo — it is SSO, legacy data, security review, and day-2 ops. The failure modes that kill production agents, and how to harden before handoff. *SSO. Legacy fields. Security review. The parts that decide whether your agent survives week two.* If you only watch launch videos, Forward Deployed Engineering looks like clever prompts and a clean UI. In the field, roughly **eighty percent of the real work** is not the demo. It is getting production credentials, surviving identity and access rules, cleaning data nobody documented, passing security review, and leaving something operators can run when you are offline. That is the job described on my Forward Deployed Engineer hub — and the reason FDE is not the same as a polished proof of concept or a ticket-only freelance milestone (FDE vs SE vs freelance). This post is the unglamorous map: what breaks, why demos lie, and what to harden before you call it shipped. The short version A Forward Deployed Engineer spends most of the engagement on **environment reality**: authentication and authorization, messy or incomplete data, network and vendor constraints, security and compliance gates, observability, and handoff. The model call is often the easy line in the stack. If you cannot pass SSO, ground answers in the customer's real corpus, and survive a security questionnaire, you do not have a product — you have a laptop demo. Why demos lie Demos are optimized for narrative: Demo world | Production world Shared admin API key | SSO, SCIM, least privilege, token rotation One clean PDF folder | Three CMSs, stale Notion, PDFs with scanned junk Happy-path prompts | Angry users, empty retrieval, policy blocks "It works on my machine" | Change windows, allowlists, VPC, vendor rate limits Ship Friday | Security review in three weeks FDE work starts when someone says: *"Cool. Now put it on our stack."* The four buckets that eat the calendar 1. Identity: SSO, roles, and who may see what **What breaks:** OAuth redirect URIs wrong per environment. Role claims that do not match who should see which docs or tickets. Service accounts that work in staging and die in prod. "Just use a shared password," which security will reject later. **What good looks like:** an explicit identity provider path. A role map written down — who can admin, who can chat, who can export. Secrets in a vault or platform store, not in a slide deck. A test user that is *not* the founder's god-mode account. **The AI-specific twist:** if the agent can retrieve customer data, identity is a **data boundary**, not a login screen. Wrong role means wrong documents in context, which is a silent leak. 2. Data: legacy fields, half-true docs, empty retrieval **What breaks:** field names that mean three things across systems. Help centers with outdated articles ranked high in search. PDFs that are scans with no text layer, sold as "knowledge base." CRM notes that are gold but never in the corpus. "We have an API" that turns out to be CSV exports on Fridays. **What good looks like:** a corpus inventory with source, owner, refresh rate and trust level. A thin slice on **live** content before any UI polish. Citation or source tracing so operators can audit answers. A plan for stale content — drop it, downrank it, or flag it. This is why citation-grounded paths matter. DotChat is built around "every claim back to a page," not freeform confidence. SupportPilot only works if the help docs are the real operational source — and escalation exists when they are not. **The FDE move:** prove retrieval quality on their mess before debating model brands. 3. Security review: the calendar you did not estimate **What breaks:** questionnaires about retention, subprocessors and log access. "No training on our data" requirements versus default vendor settings. Pen-test or vendor risk that blocks go-live after the demo wow. Audit needs — who asked what, what the model saw, what was blocked. **What good looks like:** a one-page data flow diagram (user → app → model → tools → logs). A clear retention and redaction story. Policy checks **before** the model when stakes are high. An exportable audit trail for incidents. Guardrail-shaped work looks like Mindoor, which policy-inspects before the model and exports an audit PDF, and AgentPayOps, which pairs spend policy with a decision log. Security is not a slide; it is product behaviour. **The FDE move:** start the security conversation in week one, not after UI freeze. 4. Day-2 ops: when you are not in the Slack thread **What breaks:** no logs when the agent gets weird. No owner for prompt or corpus updates. An escalation path that dumps users into a void. Rate limits and cost spikes with no alert. One engineer's laptop as the bus factor. **What good looks like:** minimal observability — request id, user and role, retrieval hits, tool calls, errors. A runbook of common failures and who fixes them. A kill switch or feature flag. A handoff session with the people who will operate it. If the system requires you forever, you did not finish the FDE loop — you rented yourself as production infrastructure. The five-step loop on the hub ends in **handoff** for a reason. A hardening checklist before you say "shipped" Use this as an acceptance gate, and steal it for proposals. **Identity:** SSO or agreed auth works for a non-admin test user. Roles documented and enforced on data access. Secrets not in client-side code or chat history. **Data:** corpus sources listed with owners. Thin slice evaluated on real content, not fixture PDFs. Empty-retrieval and low-confidence behaviour defined. Citations or source links where claims matter. **Security shape:** data flow diagram shared with their security contact. Retention, training and subprocessor answers written down. Policy or allowlist gates where required. An audit log path for sensitive actions. **Ops and handoff:** basic logging and error visibility. Escalation path tested end to end. A runbook, even one page. A named owner on their side after handoff. Cost and rate-limit awareness. If half of this is red, you have a demo, not a deployment. How this shows up in my work Project | The "80%" surface SupportPilot | Real help docs, escalation, ticket summary — not chatbot cosplay DotChat | Retrieval and citations on real PDFs; the wrong page is a failed job webflow-agent-kit | Typed, validated tools so agents do not freestyle CMS damage Mindoor | Policy before model; audit export AgentPayOps | Spend decisions logged, and blocked when policy says no Open tools and case studies are how field pain becomes reusable patterns — the feedback half of FDE, not only the embed half. How to talk about this with clients Do not lead with fear. Lead with a path. We ship a thin slice on your real stack in week one. In parallel we map SSO, data sources and security questions. Go-live means hardening and a runbook, not only a happy demo. And here is what you will own after handoff. That framing separates FDE delivery from proof-of-concept theatre and from "ticket done, environment unknown." More on those role lines in FDE vs Solutions Engineer vs freelance. Why do people say 80% of FDE work is not the demo? Because production success is gated by identity, data quality, security and operations. The model path is often a small fraction of calendar time once the customer environment is real. Should security review wait until the product is finished? No. Start data-flow and vendor questions early. Late security review is how demos die on the calendar. Is this only an enterprise problem? No. Startups hit the same issues with Google Workspace SSO, Notion-as-CMS, and auth delayed until later. Scale changes paperwork volume, not the buckets. What is the fastest way to de-risk an AI pilot? A thin vertical slice on live data, an explicit auth path, written empty-retrieval and escalation behaviour, and a one-page data flow for security. Polish the UI after those are green. ### FDE vs Solutions Engineer vs Freelance Developer https://anilpervaiz.com/blog/fde-vs-solutions-engineer-vs-freelance (Forward Deployed Engineering · 2026-08-18) Forward Deployed Engineer vs Solutions Engineer vs freelance developer — who owns production code, who embeds with operators, and which role you actually need for AI in a real customer stack. *Same customer. Three different jobs. Pick the delivery model, not the trendiest title.* People mix up **Forward Deployed Engineer (FDE)**, **Solutions Engineer (SE)**, and **freelance developer** because all three can sit on calls, write some code, and talk about AI. The difference is not vocabulary. It is **what they are paid to own**: a closed deal, a shipped milestone, or a production system operators can run without them. If you need the short role definition first, read the Forward Deployed Engineer hub. For the day-to-day loop, see what a Forward Deployed Engineer actually does. Quick answer Role | You hire them to… | Success looks like… Forward Deployed Engineer | Embed and ship production software in the customer's environment | Operators run the system; messy integrations are owned; field lessons feed product or open tools Solutions Engineer | Win and support the technical sale — demos, PoCs, architecture credibility | Deal progresses; PoC accepted; handoff to delivery or CS Freelance developer | Deliver scoped work against a brief or tickets | Milestone accepted; invoice cleared **One line:** FDE = engineer first, customer-embedded second. SE = technical sale first. Freelance = scoped delivery first. Side-by-side comparison Dimension | Forward Deployed Engineer | Solutions Engineer | Freelance developer Primary job | Ship production software in the customer's environment | Win and support the technical sale or PoC | Deliver scoped tickets or projects Writes production code? | Yes — core of the role | Sometimes light; often demos and reference architectures | Yes Embeds with operators? | Deeply — workflows, constraints, day-2 ops | During sales and onboarding | Project-scoped Owns messy integrations? | Yes — SSO, legacy data, security review, half-documented APIs | Advises; rarely owns long-term | Only if the contract says so Feeds product or open tools? | Expected — field to product or OSS | Sometimes — win stories, feature requests | Rarely Typical artifacts | Production paths, runbooks, evals, guardrails, handoff docs | Decks, demos, PoC environments, RFP answers | PRs, tickets, delivered features Time horizon | Engagement until the system is operable | Pre-sale, close, light post-sale | Sprint, fixed scope, or retainer tickets Failure mode | Demo that never hardens; hero dependency on the engineer | PoC that cannot become production | Spec met, business outcome missed Success metric | Working system operators can run | Deal closed or PoC accepted | Milestone accepted Same project, three different outcomes Imagine a company says: *"We need an AI support agent on our help docs."* Solutions Engineer path Discovery call, competitive landscape, architecture diagram. A polished demo on sample docs. A PoC that answers happy-path questions. Hand-off note: "Engineering will productionize after signature." **Win if** the deal closes. **Risk if** nobody owns SSO, escalation, evals, or the mess in the real help center. Freelance developer path Scope: "Build a chat widget and RAG over docs." Ships against the written brief. Done when acceptance criteria on the ticket pass. **Win if** the milestone is fair and complete. **Risk if** the brief never included escalation, citation quality, or operator handoff — and nobody is paid to discover that. Forward Deployed Engineer path Scope the real outcome: deflect tier-1, escalate cleanly, attach ticket summaries. Map *their* docs, auth, ticketing tool, and who sits on the queue. Thin vertical slice on **live** help content. Harden: grounding, escalation rules, logging, simple evals. Handoff: a runbook so support leads can operate it. That is the shape behind work like SupportPilot (docs RAG with human escalation) and DotChat (answers cited to source pages) — not "call an API and ship a widget." Where the titles blur, and how to unblur them "Our SE writes a lot of code" Some Solutions Engineers are deeply technical. The test is still: **are they measured on production ownership in the customer environment, or on pipeline and PoC?** If the answer is pipeline, it is SE — even if they open a PR. "Our freelancer embeds with us" Some freelancers work like FDEs: on-site energy, full systems thinking, leftover runbooks. The test is the **contract and incentives**. If scope is tickets and "done" is acceptance of a milestone with no field feedback loop, it is freelance delivery — even if the person is excellent. "FDE is just a fancy consultant" Consultants often optimize for recommendations and workshops. FDEs optimize for **running software**. Overlap exists; the artifact differs. Prefer people who leave systems and docs, not only slides. "In AI startups everyone is a bit of everything" True early on. Titles still matter when you hire, price, or explain yourself. If you sell FDE and deliver SE demos, trust dies. If you sell freelance tickets and the client expected embedded production ownership, the engagement explodes mid-way. Which one do you need? **Hire or engage an FDE when** the idea is stuck in demo mode; the environment is messy (legacy CMS, many tools, compliance, incomplete APIs); you need someone who can talk to founders **and** merge production code; you are embedding agents into support, CMS or Webflow ops, internal tools, or spend workflows; or you want leverage left behind — runbooks, patterns, open tools — rather than a black-box exit. **Hire a Solutions Engineer when** you are selling a platform and need technical credibility in the sales cycle, PoCs and architecture narratives unblock deals, and delivery after signature is owned by another team. **Hire a freelance developer when** scope is clear, acceptance criteria are honest, and the environment is already understood; you need capacity on a defined surface; and you do not need embedded discovery or long-lived production ownership from that person. **Hybrid reality:** many companies need SE **and** FDE. Many founders need a freelancer who can *temporarily* operate in FDE mode — price and scope that explicitly, or you will underpay and over-expect. How I use these labels for myself I position as a **Forward Deployed Engineer** for AI agents, RAG, and automation inside real stacks (Webflow, Supabase, n8n, Next.js): scope with the client, ship against their data, harden with guardrails, hand off something that survives without me. That is not the same product as taking tickets indefinitely, and not the same as running an enterprise sales PoC. Proof lives in case studies and open source — SupportPilot, DotChat, webflow-agent-kit, AgentPayOps — and the role definition lives on the Forward Deployed Engineer page. If you are hiring for FDE-shaped work: get in touch, or take the CV from the hub page. Decision checklist you can steal Before you post a role or sign a contract, answer in writing: 1. What is the **success metric** — deal, milestone, or operable system? 2. Who **owns production** after week four? 3. Is discovery **in scope**, or is the brief already true? 4. Do we need **field feedback** into a product or open toolkit? 5. What does **handoff** look like if this person disappears for two weeks? If (1) is an operable system, (2) is unclear, (3) is yes and (4) is yes — you are describing an FDE. Write the title and the contract to match. Is a Forward Deployed Engineer just a Solutions Engineer who codes more? No. Coding volume is not the divider. Ownership of production outcomes in the customer environment is. Solutions Engineers can code; FDEs are measured on systems that run after the meeting ends. Can a freelancer work as an FDE? Yes, if the engagement is scoped that way: embed, real data, harden, handoff, and optionally feedback into tools or product. Call it what it is and price the discovery and production risk — do not hide FDE work under a cheap ticket rate. Which role is best for AI agent projects? If the agent must live in a customer's tools, data and ops, you need FDE-shaped delivery whatever the badge says. If you are selling a platform and need a credible proof of concept in the sales cycle, you need a Solutions Engineer. If the architecture is already decided and you need hands, freelance capacity can be enough. Where should I start if I am still confused? Read the Forward Deployed Engineer hub, then the field guide, what a Forward Deployed Engineer actually does. ### What a Forward Deployed Engineer actually does https://anilpervaiz.com/blog/what-a-forward-deployed-engineer-actually-does (Forward Deployed Engineering · 2026-08-12) A Forward Deployed Engineer embeds with customers to ship production software in messy real stacks — not demos. The day shape, the five-step loop, and how AI FDEs differ from freelancers and solutions engineers. *Not the job description. The real loop: scope, build in their stack, harden, hand off — and feed the field back into product.* A Forward Deployed Engineer (FDE) is a software engineer who embeds with the customer to turn messy business problems into production software. The job is not "write features from a ticket queue." It is discovery in a live environment, code against real data and constraints, systems that operators can run after you leave, and patterns that feed back into the product or open tools. If you want the short definition and how I position the role, start on my Forward Deployed Engineer page. This post is the longer field version. The job in one paragraph FDEs sit between the customer's reality and the company's (or your own) platform. You meet operators, map how work actually happens, design something that fits *their* stack, write production code, debug integrations that only break on real SSO and legacy data, ship with guardrails, document the handoff, and report what the core product should absorb next. At places like Palantir the loop is classic: meet customers, design and test software, configure the product, communicate feedback home. In AI companies the same loop shows up as agents, RAG, evals, and workflow automation inside the customer's tools — not a sandbox demo. What an FDE does not do Not the job | Why people confuse it Pure sales engineering | SE optimizes for the deal and PoC; FDE optimizes for a system operators run Ticket-only freelance | Freelance can ship code; FDE owns embedding, messy integrations, and feedback loops Research science | You may use models; you are measured on production outcomes, not papers Remote ticket hero with no customer contact | The "forward" part means you are in the customer's context Full-stack is a **skill set**. FDE is a **delivery model**. I break the comparison down on the FDE hub. A realistic day, not a fantasy calendar No two days are identical. Most FDE weeks mix roughly 40% building against the customer's real system, 30% client syncs, triage and live debugging, 20% internal alignment and pattern reporting, and 10% documentation and handoff. Morning — triage and truth Start with what broke overnight or what changed in their environment. Auth rotated. A CMS field renamed. A webhook silently failed. An agent answered confidently and wrong. You re-establish ground truth before you write new code. Midday — build on live paths This is the core block. You are not polishing a slide deck. You are wiring retrieval to *their* docs, putting policy checks in front of the model, fixing the n8n path that only fails on production credentials, or shipping a thin vertical slice that proves the risky integration first. Afternoon — align and leave leverage Sync with stakeholders. Capture what should become a reusable pattern. Write the short runbook so the next person is not blocked when you step away. If you work product-side, this is also when you push field feedback home. If you are independent, this is when you harden open tools and case notes. That rhythm is why the role feels closer to embedded product engineering than to classic consulting decks. The five-step loop I actually run This is the same loop on my FDE page. It is how I keep speed without abandoning production quality. 1. Scope Clarify the business outcome, constraints, owners, and how we will judge success. "We need AI support" is not a scope. "Deflect tier-1 questions from help docs, escalate cleanly to a human, and attach a ticket summary" is a scope. 2. Map the stack Data sources, tools (Webflow, Supabase, n8n, CRMs), auth, who operates the system day to day. Most failures are stack and ownership failures, not model failures. 3. Thin vertical slice One end-to-end path on real data. Prove the risky integration first. A beautiful UI on fake data is not an FDE win. 4. Harden Validation, guardrails, logging, escalation, simple evals, failure modes. For AI work that often means citation checks, spend or policy gates, and "what happens when the model is wrong." 5. Handoff Docs, enablement, reusable patterns. The system should survive without you. If everything depends on you being in Slack, you did not finish the job. What this looks like in AI work AI made the FDE model louder because models do not just plug in. Someone has to own integration, evals, and operations inside the customer's world. Support agents that escalate like adults SupportPilot is the pattern: an embeddable agent trained on client help docs, with clean human escalation and ticket summaries. The FDE work is not "call an LLM." It is grounding, escalation rules, and a handoff operators trust. RAG that cites sources DotChat is PDF chat where every claim points back to a source page — retrieval and answer split so you get grounded responses. The FDE work is pipeline design, evaluating "did we cite the right page," and failure modes when retrieval is empty. Agent tools that are safe on a real CMS webflow-agent-kit is TypeScript-first, Zod-validated tooling so agents touch Webflow with typed, safer operations — not raw hope that the CMS script works. Open source is also the feedback loop: field patterns become reusable tools. Guardrails before the model spends or answers AgentPayOps puts policy and audit on autonomous spend. Mindoor inspects clinic messages before they hit the model and exports audit-ready trails. That is classic FDE territory: production constraints, compliance shape, and systems that survive scrutiny. Multi-agent systems under pressure Autopsy runs multi-agent research and debate to a forensic verdict. The lesson for FDE work is orchestration and synthesis, not a single prompt demo. The hard parts nobody puts in the LinkedIn post Roughly eighty percent of the real job is not the demo. It is SSO, roles, and who is allowed to see which data. Legacy fields and half-documented APIs. Security review and change windows. Operators who will ignore your tool if it adds three clicks. Evals that catch confident nonsense before customers do. Writing the boring runbook so the system does not depend on your memory. If you only enjoy greenfield prototypes with perfect fixtures, FDE will frustrate you. If you like making something work in a living environment, it is one of the highest-leverage engineering jobs in AI right now. Skills that actually matter **Technical:** strong generalist coding (TypeScript, Python), APIs, auth, data plumbing, enough LLM and RAG literacy to ship and evaluate, automation with n8n or Zapier when the customer lives there. **Delivery:** scoping ambiguous problems, stakeholder communication, writing for handoff, saying no to scope that cannot survive production. **Judgment:** when to configure versus build, when the model is the wrong tool, when to push a pattern upstream or into open source. I keep a tighter capability list on the FDE hub. When a company actually needs an FDE Engage this profile when the idea is stuck in demo mode, the environment is messy (legacy CMS, many tools, compliance, incomplete APIs), you need someone who can talk to founders *and* merge production code, you are embedding agents into support, CMS ops, internal tools or spend flows, or you want documentation and leverage left behind rather than a black-box contractor exit. If you only need a landing page or a one-off script, you may not need an FDE. If you need production AI inside a real stack, you do. How I work this role independently I am based in Dubai and ship as a customer-facing engineer: scope with the client, build agents, RAG, and automation against their data and systems (Webflow, Supabase, n8n, Next.js), then hand over documented systems with guardrails. The public proof lives in case studies and open source — and the role definition lives on the Forward Deployed Engineer page. If you want the CV-shaped version for hiring loops, download it from that page. If you want to talk about an engagement, get in touch. What does a Forward Deployed Engineer do day to day? A Forward Deployed Engineer scopes real customer problems, writes production code in the customer's environment, debugs integrations across data, auth and APIs, ships systems operators can run, and documents handoff so the work survives after the engagement. Is FDE the same as a full-stack developer? No. Full-stack is a skill set. FDE is a delivery model: embed with the customer, own outcomes through production, and close the loop with field feedback. Is FDE the same as a Solutions Engineer? Usually no. Solutions Engineers often optimize for the technical sale and proof of concept. FDEs optimize for a production system in the customer's stack. Do FDEs only work at big AI labs? No. The title is common there, but the model applies anywhere complex software — especially AI agents and automations — must land in a real customer environment, including startups and agencies. What should I read next? Start with the hub, Forward Deployed Engineer. Then the case studies under Work. ### AI agency vs. in-house vs. fractional: how to staff your AI work https://anilpervaiz.com/blog/ai-agency-vs-in-house (AI Architecture · 2026-05-26) The real trade-offs between hiring an AI agency, building an in-house team, and bringing in a fractional AI lead — and which fits your stage. You've decided AI is worth real investment. Now the harder question: who builds it? There are three honest options, and the right one depends almost entirely on your stage. Option 1: An AI agency Agencies start fast and carry broad experience, but you pay for overhead you don't benefit from, and the knowledge often leaves when the engagement ends. They fit a one-off build with a clear spec — a defined feature, shipped and handed over. They're a poor fit when AI is becoming core to your product and you need someone who learns your domain deeply over time. Option 2: Hire in-house A full-time AI engineer is the right end state once AI is a core surface and you have enough work to keep them busy. The catch is timing and cost: senior AI talent is expensive and slow to hire, and a single hire takes months to get productive in your codebase. Hiring in-house before you've validated what to build is how teams end up with an expensive person and no shipped feature. Option 3: A fractional AI lead The middle path, and usually the right one for SMBs and agencies: someone senior who works part-time as your AI lead — architecting the system, prototyping the riskiest pieces, shipping the first version, and helping you hire in-house when volume justifies it. You get senior judgment without a full-time salary, and the knowledge stays documented in your codebase instead of walking out the door. (For what that role does day to day, see the shape of an AI architect.) How to choose, by stage - **Exploring** (nothing shipped yet): start with a short audit plus one fractional engagement. Validate before you hire. - **One clear build**: an agency or fixed-scope project works — just insist you own the code. - **AI becoming core, low volume**: a fractional AI lead bridges you until a full-time hire is justified. - **AI core, high volume**: build in-house, ideally with a fractional lead helping you hire and onboard. The cost question Roughly: a fixed-scope build runs a few thousand to low five figures, a fractional lead is $5k–$15k/month part-time, and a senior full-time hire is a six-figure salary plus ramp time. I broke the numbers down in what an AI consultant costs. The cheapest option is rarely the one that ships the right thing fastest. The takeaway Don't staff for the company you'll be in two years; staff for the next shipped feature. For most teams that means an audit, then a fractional lead, then in-house once the work is proven. That fractional-lead path is exactly how I work with agencies and SMBs — see services or book a call. ### How to add AI to your SaaS (without a rebuild) https://anilpervaiz.com/blog/how-to-add-ai-to-your-saas (AI Architecture · 2026-05-26) A practical sequence for shipping your first real AI feature into an existing product — what to build first, what to skip, and how not to break what already works. Most SaaS teams don't need an AI strategy. They need one AI feature that earns its place, shipped without breaking the product that already pays the bills. Here is the sequence I use to get there. Start with one workflow, not "AI" "Add AI" is not a project. Find the single workflow where AI moves a number you already track: support response time, onboarding length, time-to-first-value, hours lost to a manual task. Pick one. Teams that try to "become an AI company" in a quarter ship nothing; teams that fix one painful workflow ship in two weeks and learn what to do next. Where AI actually fits in a SaaS In practice the first feature is almost always one of these: - A support or docs assistant that deflects repetitive tickets. - A drafting step that turns a blank field into a starting point (replies, summaries, descriptions). - A retrieval layer so users can ask questions across their own data. - A behind-the-scenes classification or routing task no user ever sees. None of these is "a chatbot bolted to the homepage." The best AI features disappear into a workflow users already have. Build vs. buy, the honest version Before building anything, check whether an existing tool does 80% of the job. If one fits, use it and spend the budget elsewhere. Build custom only when the feature touches your proprietary data or your core differentiation. I've talked clients out of builds more than once — it's the fastest way to earn trust. (More on that call in what an AI architect does.) The 14-day path to your first feature - Days 1–2: pick the workflow and the metric; define what "good enough to ship" means. - Days 3–6: prototype the riskiest piece (usually retrieval quality or the prompt) behind a flag, with a small eval set so you can measure it. - Days 7–11: wire it into the real product surface, with guardrails and a human fallback for low-confidence cases. - Days 12–14: instrument it, ship to a slice of users, and watch the metric. The point isn't speed for its own sake. A live feature in front of real users teaches you more in two weeks than a quarter of planning. What not to do - Don't fine-tune when you mean retrieval. If the model needs to know your facts, that's a retrieval problem, not a training one. - Don't ship without evals. "It felt better in the demo" is how AI features quietly regress. - Don't make it un-ownable. If only an outside contractor can keep it alive, you bought a liability. - Don't let it touch money, health, or legal data without guardrails and a human in the loop. The takeaway Adding AI to a SaaS is not a rebuild. It's picking one workflow, proving it with a metric, and shipping the smallest version that works. Do that once and the second feature is obvious. If you want help picking the first one, that's exactly what the AI Audit & Roadmap is for. Book a 15-minute call and we'll find your highest-leverage workflow. ### What does an AI consultant cost in 2026? https://anilpervaiz.com/blog/ai-consultant-cost (AI Architecture · 2026-05-23) Real 2026 pricing for AI audits, builds, retainers, and fractional leads — what drives the number, what each tier should cost, and how to avoid overpaying. Most independent AI consultants charge $500 to $5,000 for an audit, $2,000 to $25,000 for a project build, and $1,500 to $8,000 a month on retainer. Agencies charge two to four times that for the same work. Below is what sits behind each number, and how to tell a fair quote from a padded one. How much does an AI consultant cost in 2026? Independent AI consultants and architects fall into four pricing tiers, in USD. An audit runs $500 to $5,000, a fixed-scope project build $2,000 to $25,000 or more, a monthly retainer $1,500 to $8,000, and a fractional AI lead $5,000 to $15,000 a month. - Audit or roadmap: $500 to $5,000 for a fixed-scope review - Project build: $2,000 to $25,000+ depending on complexity - Monthly retainer: $1,500 to $8,000 a month for ongoing work - Fractional AI lead: $5,000 to $15,000 a month for embedded, part-time leadership Agencies and larger consultancies charge two to four times these numbers for the same work, mostly to cover overhead you do not benefit from. What makes one quote higher than another? Three things move the number more than anything else: how clearly the work is scoped, how ready your data is, and how expensive a wrong answer would be. Everything else is noise. - Scope clarity. A vague brief is expensive because someone has to absorb the risk of the unknown. The tighter your spec, the lower the quote. - Data readiness. If your data is clean and accessible, a build is fast. If it is scattered across five tools in three formats, half the budget goes to plumbing before any AI happens. - Failure tolerance. A marketing chatbot that is occasionally wrong is cheap. A system that touches money, health, or legal data needs guardrails, evals, and audit trails, and that is where cost climbs. What should an AI audit cost? Between $500 and $5,000. That buys a few days of work ending in a document: where AI helps, what to build first, rough cost, and what to skip. It is the highest-leverage money you will spend, because it stops you funding the wrong thing. Mine is fixed at $500 precisely so it is a no-brainer. What should an AI project build cost? Between $2,000 and $25,000 for one well-scoped feature shipped to production. A focused chatbot or RAG layer lands near the bottom of that range; a multi-step agent with integrations and a real eval harness lands higher. Fixed scope and fixed price should be the default. If someone quotes hourly for a well-defined build, that is risk transferred from them to you. Hourly, fixed-fee, or retainer — which should you choose? Use fixed-fee when the outcome is clear, a retainer when the work is continuous, and hourly almost never. Hourly only makes sense for genuinely open-ended discovery, and even then a capped audit is usually the better instrument. A retainer at $1,500 to $8,000 a month covers ongoing iteration: new flows, prompt tuning, evals, monitoring, and fixes. It earns its keep once AI is live and you need it to keep improving without a procurement cycle every time. What does a fractional AI lead cost? $5,000 to $15,000 a month for part-time senior leadership: architecture ownership, hiring help, and hands-on building, usually 10 to 15 hours a week. It is the right move for teams making AI a core surface but not ready for a full-time hire. How do you avoid overpaying for AI consulting? Start with an audit, insist on fixed scope, and refuse any engagement that cannot tell you how success will be measured. Those three rules remove most of the ways this goes wrong. - Start with an audit before any build. It pays for itself. - Insist on fixed scope and fixed price for builds. - Ask how they will measure success. No eval, no deal. - Make sure you own the code. Everything should be handed over clean. - Skip the build if an off-the-shelf tool does 80% of it. A good consultant says so. Is hiring an AI consultant worth it? Yes, if the project saves or earns more than it costs within a year. A $5,000 build that removes 20 hours a week of manual work pays for itself in about a month. A $5,000 build that "adds AI" with no metric attached is a $5,000 science project. For a straight answer on what your specific project should cost, send a brief or book a 15-minute call. I give you a range on the call, not after three meetings. You can also compare the four engagement tiers on the services page. ### What is an AI architect (and when you actually need one)? https://anilpervaiz.com/blog/what-is-an-ai-architect (AI Architecture · 2026-05-20) The honest definition of an AI architect, how the role differs from an AI or ML engineer, and a straight answer on whether you need to hire one. "AI architect" is a title that barely existed five years ago, and half the people using it mean different things. Here is the working definition I use, plus an honest answer to the question most founders are really asking: do you need to hire one, or can your existing team handle it? What does an AI architect actually do? An AI architect designs the whole system around a model, not just the prompt. The prompt is maybe 10% of a working AI feature. The other 90% is the part nobody demos: where the data lives, how the model retrieves it, what happens when it is wrong, how you measure whether it is improving, and how a non-AI engineer maintains it six months later. On day one of a project I am making decisions like: - Which model, and why (cost, latency, and accuracy trade-offs) - Whether this needs retrieval, fine-tuning, or neither - Where the guardrails live and what the fallback is when confidence is low - How we will evaluate quality before and after launch - What the human-in-the-loop checkpoints are - How it integrates with the systems you already run A prompt engineer optimizes wording. An AI architect makes sure the wording is the smallest, last problem you have. AI architect vs. AI engineer vs. ML engineer These get used interchangeably, so quickly: - An ML engineer trains and deploys models. They care about weights, datasets, and GPUs. - An AI engineer builds applications on top of existing models. They care about APIs, latency, and product behavior. - An AI architect sits a level up: they decide what to build, which approach fits the business, and how the pieces connect, then often build the first version themselves. Most startups in 2026 do not need to train models. They need someone who can take an off-the-shelf model and turn it into a feature that survives real users. That is the architect's job. When do you actually need one? You probably need an AI architect when: - You shipped a demo that wowed everyone and quietly fell apart in production. - Your team can call an API but is not sure how to tell whether the output is good. - You are about to spend real money and want someone to tell you what not to build. - AI is becoming a core surface of your product, not a side feature. You probably do not need one when an off-the-shelf tool already does 80% of the job. A good architect tells you that on the first call instead of selling you a build. I turn down work for this exact reason more often than you would expect. What does "good" look like? The tell is whether someone brings receipts. Ask a candidate how they would know their AI feature works. If the answer is "it will feel better," keep looking. If the answer involves an eval set, a baseline, and a number they are trying to move, you are talking to an architect. The other tell is restraint. The best AI work in 2026 is boring on purpose: observable, well-tested, and built so the model is a component you can swap, not a black box your business depends on. (Google says the same thing about content, by the way: it rewards work that demonstrates real experience and expertise, not volume.) How to start without overcommitting You do not have to hire a full-time AI lead to find out if this is worth it. The lowest-risk first step is a short audit: someone spends a few days mapping where AI moves a real number for you and what it takes to ship, then hands you a plan you own either way. That is exactly the AI Audit & Roadmap I run, and it is how most of my engagements start. To see the kind of systems this produces, the AI Lab has live demos you can poke at, and recent work shows the metric each one moved. If you are weighing whether to hire, book a 15-minute call and I will tell you honestly whether you need an architect or just a better tool. ### The shape of an AI architect https://anilpervaiz.com/blog/the-shape-of-an-ai-architect (AI Architecture · 2026-04-22) What separates AI-native builders from prompt tinkerers — and the seven decisions I make in the first days of almost every project. An AI architect doesn't just write prompts. They model the whole stack: where the data lives, how retrieval happens, what the model decides, and where humans stay in the loop. That is the difference between a demo that wins a meeting and a system that survives real users. If you are still deciding whether you need this role at all, I wrote a separate piece on what an AI architect is and when you actually need one. This one is about the work itself: the seven decisions I make in the first days of almost every project, before a line of feature code gets written. 1. What is the one job to be done? Most AI projects fail because they try to be impressive instead of useful. Before anything else, I find the single workflow where AI moves a number the business already tracks. One job, one metric. A support team drowning in tickets. A sales team losing hours to research. A content pipeline stuck in review. Everything else waits until that one thing works. 2. Build, buy, or skip? The most valuable sentence an architect says early is "you don't need to build this." If an off-the-shelf tool does 80% of the job, we use it and spend the budget where it actually matters. Custom only earns its keep when it moves a metric a tool can't reach. Saying this out loud has cost me projects. It has also earned me every long-term client I have. 3. Retrieval, fine-tuning, or neither? Teams reach for fine-tuning when they mean retrieval. If the model needs to know your facts, that is almost always retrieval, not training. Fine-tuning changes behavior and tone, not knowledge, and it is expensive to maintain. Most products in 2026 need good retrieval and a sharp prompt, nothing more. 4. Where do the guardrails live? A model will eventually produce something wrong, off-brand, or unsafe. The decision is not whether that happens but where you catch it: input validation, output checks, a moderation pass, or a human gate. I decide this on day one, because retrofitting guardrails into a shipped system is painful and usually means a rebuild. 5. How will we know it works? If you can't measure quality, you can't improve it, and you certainly can't trust it. I build a small eval set early: real inputs, expected behavior, and a baseline number. Then every change is measured against it. "It feels better" is not a launch criterion, and any architect who offers it as one is guessing. 6. Where do humans stay in the loop? The best AI systems are honest about uncertainty. They cite sources, surface confidence, and escalate to a person at the right threshold. That handoff is a design decision, not an afterthought. Get it right and users trust the system even when the model is unsure. 7. Who runs this after launch? An AI feature is not done when it ships. Someone has to monitor it, watch cost, and catch drift as the world changes around it. I design for that from the start with clean code, observability, and a runbook, so the team owns it without me. If only the architect can keep it alive, the architecture failed. The thread that connects all seven None of these decisions are about the model. They are about the system around it. That is the whole job, and it is why the same prompt can be a toy in one product and load-bearing infrastructure in another. This is exactly the thinking behind the AI Audit & Roadmap I run at the start of most engagements. You can poke at the systems it produces in the AI Lab, or book a 15-minute call and tell me about the one job you would want AI to do. ### Shipping Claude Code for real work https://anilpervaiz.com/blog/shipping-claude-code-for-real-work (Engineering · 2026-03-14) A field report from running Claude Code on production codebases — the patterns that scale, the failure modes that look like success, and the rituals I keep. After six months of running Claude Code as a daily driver across three production codebases, I stopped reaching for autocomplete-style copilots. The difference isn't the model. It's how you work with it. Here is the field report: the prompt patterns that scale, the failure mode that looks like success, and the small infrastructure investments that turn a clever assistant into an actual teammate. Treat it like a contractor, not autocomplete Autocomplete finishes your line. An agent does a task. The mental shift is to stop thinking in keystrokes and start thinking in work orders: "add a rate limiter to these three routes, match the existing error shape, and show me the diff." The closer your request looks to a ticket you would hand a competent contractor, the better the result. The prompt patterns that actually scale Three habits did most of the work: - Point at examples. "Match the pattern in this file" beats any amount of description. - Constrain the surface. Tell it which files to touch and which to leave alone. Unbounded tasks produce unbounded diffs. - Ask for the plan first on anything non-trivial. A 30-second plan review catches the wrong approach before it writes 300 lines. The failure mode that looks like success The dangerous output isn't the obviously broken one. It's the confident, plausible diff that passes a glance and quietly does the wrong thing: a test that asserts nothing, an error path swallowed, a security check skipped because it "wasn't in scope." This is the same lesson as building agents that don't hallucinate: the model is most dangerous when it is wrong and smooth about it. The fix is simple and non-negotiable — never merge what you didn't read. Review is the new bottleneck When the agent writes faster than you can read, your review becomes the constraint. I leaned into it: smaller diffs, clearer commits, and a habit of asking "what would make this wrong?" before merging. The teams that get burned by AI coding tools are the ones that let velocity outrun review. The infrastructure that turns it into a teammate A few cheap investments compounded: - A tight project doc stating conventions, so I'm not re-explaining the stack every session. - Fast tests and a type-checker it can run itself, so it gets a feedback loop instead of guessing. - Small, reviewable commits, so when something is wrong I can see exactly where. None of this is exotic. It is the same hygiene that makes human teams fast — the agent just rewards it more obviously. Where I don't use it I still write the hard parts myself: the security-sensitive code, the gnarly state machine, the thing where being 95% right is worse than not shipping. The agent is fast on the well-defined 80%. The remaining 20% is where the judgment lives, and that judgment is exactly what clients pay an AI architect for. Claude Code didn't replace engineering judgment. It moved where I spend it: less typing, more reviewing and deciding. If you want this kind of velocity wired into your own stack, the AI Lab shows what I've shipped, and you can book a call to talk about it. ### Websites that grow with the brand https://anilpervaiz.com/blog/websites-that-grow-with-the-brand (Design · 2026-02-02) The modular system that lets non-technical teams ship pages without breaking the design — and why that discipline is what makes the AI I build ship clean. Most marketing sites collapse the moment a non-designer touches them. A founder edits one headline and the layout breaks. Someone pastes a long testimonial and the grid blows out. Within a month, the site nobody wanted to maintain becomes the site nobody can. I have shipped more than 40 production sites, and the ones that survive a year of edits all share the same backbone. That same discipline, it turns out, is exactly what makes the AI I build today ship clean: constrain the inputs, define the defaults, and be clear about what is allowed to change. Why a prettier CMS doesn't fix it The problem isn't the editor, it's that the system trusts the editor with too much. Free text everywhere, no constraints, no shape rules. That is the same mistake teams make with AI prompts: unbounded input, unbounded failure. The fix: a real component system The sites that last are built from a small set of components with constrained slots: - Fixed content shapes. A "feature" has an icon, a title under a character limit, and a body. There is no way to make it something else. - Sensible defaults. Empty states look intentional, not broken. - Composition, not free-form. Editors assemble approved blocks; they don't invent layouts. This is boring on purpose, and boring is the point. Content-shape rules, concretely A few rules I bake into every build: every image has a defined aspect ratio, so nothing ever stretches. Every heading has a character budget the CMS enforces. Every section has a maximum and minimum number of items. The editor literally cannot create the broken state. Designers stop policing the site and start improving it. What the constraints buy you When the system is constrained, a non-technical team ships pages every week without a designer in the loop and without anything breaking. They out-learn the team waiting on engineering tickets, and the brand stays intact. Speed and safety, which usually trade off against each other, stop fighting. The connection to AI work Here is why this matters beyond websites. The instinct that makes a good content system — constrain the inputs, design the defaults, define the failure modes — is the same instinct that makes a good AI system. A decade of shipping editable sites is why the AI features I architect don't fall apart the first time a real user does something unexpected. (For when Webflow is the right call versus code, see why I still reach for Webflow.) If you want a site your team can actually run — or AI built with the same discipline — see recent work or book a call. ### From design to deploy in 48 hours https://anilpervaiz.com/blog/from-design-to-deploy-in-48-hours (Workflow · 2026-01-15) The 48-hour sprint I run with founders to ship a high-fidelity landing page without cutting corners on accessibility or performance. Tight deadlines force you to strip away everything that doesn't matter. The trick is knowing what actually matters. Here is the 48-hour sprint I run with early-stage founders to ship a high-fidelity landing page without cutting corners on accessibility or performance. The secret isn't working faster. It's ruthless sequencing, so no hour is wasted waiting on a decision that should have been made earlier. Hours 0–6: content architecture No design, no code. We decide what the page says and in what order: the one promise, the proof, the objections, the call to action. Most "design" problems are actually unresolved content problems. Settle the words first and the layout almost designs itself. Hours 6–18: design on a locked token system I design in Figma against a fixed set of tokens — type scale, spacing, color, radius — set on hour zero. Locked tokens mean every screen is consistent by construction and translates to code without a translation layer. This is the same modular discipline that keeps sites alive long after launch. The component inventory I start with I keep a standing kit: a hero, a logo bar, a feature grid, a testimonial block, a pricing table, an FAQ, and a footer CTA. Seven blocks cover 90% of landing pages. Starting from a proven inventory instead of a blank canvas is most of where the 48 hours comes from. Hours 18–36: build with pre-wired components Build happens in Next.js, or Webflow depending on who owns the site after launch, using components already wired to the tokens. Because the design used the same token system, build is assembly, not reinterpretation. Hours 36–48: performance, accessibility, launch The last block is non-negotiable quality: image optimization, semantic HTML, focus states, contrast, and a Lighthouse pass. I keep three budgets I never break — LCP under 2.5 seconds, CLS under 0.1, and a real accessibility pass. A deadline is not an excuse to ship something inaccessible. Why it works No all-nighters. The 48 hours work because the decisions are sequenced so each phase has everything it needs the moment it starts: content before design, tokens before screens, components before build, quality as a fixed final block. That sequencing discipline is the same thing I bring to AI builds — prototype the riskiest piece first, measure, then expand. If you have a deadline that feels impossible, send me the brief. These are my favorite projects, and recent work shows how they turn out. ### Why I still reach for Webflow https://anilpervaiz.com/blog/why-i-still-reach-for-webflow (Design · 2025-12-03) Webflow, Next.js + Sanity, or AI-generated full-stack? The decision isn't about technical purity — it's about who owns the site six months after launch. I've shipped production React for ten years. I still recommend Webflow to half my clients. The reason isn't nostalgia — it's operational speed, and a hard question most teams skip: who owns this site six months after launch? The question that actually decides the stack Not "which is more powerful" or "which is more modern." The question is who edits the site once I'm gone. A marketing team that can publish pages, edit copy, and A/B test without a deploy pipeline will out-learn a team waiting on engineering tickets every single time. The best stack is the one your team can actually operate. Webflow: when the marketing team owns it If a non-technical team needs to ship and iterate fast, Webflow wins. Visual editing, no deploy pipeline, and real performance and accessibility if you build it right. The trap is treating it as a toy — a sloppy Webflow build rots as fast as any other. A disciplined component system is what makes it last. Next.js + Sanity: when engineering owns it When the site is deeply custom, integrates with a product, or needs logic Webflow can't express, I reach for Next.js and a headless CMS. More power and control, and more responsibility — someone has to maintain a codebase and a deploy pipeline. Worth it when the site is a product surface, not just marketing. AI-generated full-stack: when speed beats everything For throwaway experiments, internal tools, and rapid prototypes, AI-generated full-stack code is now genuinely fast. I use it where being 80% right today beats perfect next week, and I do not use it where it becomes load-bearing without review — the same caution I apply to shipping AI code generally. The decision matrix - Non-technical team, marketing site, frequent edits → Webflow - Custom logic, product integration, engineering team → Next.js + Sanity - Prototype, internal tool, speed over longevity → AI-generated The mistake I see most Teams pick a stack for status, not fit. A five-person startup adopts a complex headless setup because it sounds serious, then ships nothing for two months because every copy change needs a developer. The unglamorous truth: the right stack is whichever one gets your team shipping and learning fastest. Why this matters for AI work People expect an "AI architect" to push the newest stack on everything. The opposite is true. Knowing when not to reach for the powerful option is the entire skill, in stacks and in AI itself. The tool should fit the team, not the other way around. Want help picking the right stack for your situation? Book a call or see recent work. ### Building AI agents that don't hallucinate https://anilpervaiz.com/blog/building-ai-agents-that-dont-hallucinate (AI Architecture · 2025-11-18) Retrieval, guardrails, and human-in-the-loop patterns — plus three architectures I've shipped that stay grounded even when the model is unsure. Hallucination is the wrong framing. The real question is: what does your agent do when it's uncertain? A model that is confidently wrong 5% of the time isn't a bug to patch — it's a design problem to engineer around. Production AI agents fail gracefully. They cite sources, expose confidence, and escalate to humans at the right threshold — not as an afterthought, but as a core pattern. Here are the three layers that make that possible, and three architectures I've shipped using them. Layer 1: retrieval, so the model isn't guessing Most "hallucination" is the model answering from memory when it should be answering from your data. Retrieval (RAG) fixes the root cause: pull the relevant facts first, then ask the model to answer only from them. The quality of an AI feature is usually decided by retrieval quality, not the model — it's the single highest-leverage decision in the shape of an AI architect. Layer 2: guardrails, so wrong answers don't escape Retrieval reduces errors; it doesn't eliminate them. Guardrails catch what slips through: output validation, "answer only from context" instructions, citation requirements, and a refusal path when confidence is low. A system that can say "I don't know" is more trustworthy than one that always answers. Layer 3: human-in-the-loop, so stakes match oversight The higher the stakes, the more a human belongs in the loop. The design decision is the threshold: when does the agent act alone, and when does it ask? Get that line right and you get speed where it's safe and caution where it counts. Architecture 1: citation-grounded RAG (legal research) Every answer links to its source passages. If the system can't cite, it doesn't answer. Lawyers trust it because they verify in one click, and the citations make wrong answers obvious instead of dangerous. Architecture 2: multi-step approval agent (content publishing) The agent drafts and proposes, but a human approves before anything goes live. It moves work forward without ever taking an irreversible action on its own. Velocity with a safety rail. Architecture 3: real-time support bot with clean handoff When the bot hits its confidence threshold or a sensitive topic, it hands off to a human with the full conversation and its best guess attached. The user never repeats themselves, and the handoff feels like an upgrade, not a failure. You can try this class of system in the AI Lab. The pattern under all three None of these "solve" hallucination. They make the system honest about uncertainty and safe when it's wrong. That is the difference between an AI demo and AI you can put in front of customers, and it's exactly what an AI audit is for. Building something where being wrong has real consequences? Book a call. --- ## Case studies (full text) ### AgentPayOps https://anilpervaiz.com/work/agent-pay-ops Control plane for autonomous AI spend. An invoice agent attempts a paid vendor-risk report, hits an X402 challenge, the policy engine decides approve/escalate/block, and every reasoning step is logged for finance teams. The problem Autonomous agents that can spend money are powerful and terrifying. Finance teams need a control plane before they'll let an agent touch a payment. What I built A control plane for autonomous AI spend. An invoice agent attempts a paid vendor-risk report, hits an X402 payment challenge, and a policy engine decides approve, escalate, or block — logging every reasoning step for finance to audit later. The result Three clean outcomes (approve / escalate / block), programmable payments via X402, and an audit trail on every decision. It's the pattern that makes agentic payments safe enough to actually deploy. This is the guardrail-and-human-in-the-loop thinking I bring to every AI build. See more live agents in the AI Lab. ### Autopsy https://anilpervaiz.com/work/autopsy Forensic intelligence for failed companies. Twenty-four specialized AI agents across four modes — Postmortem, Pre-Mortem, Founder Mode, Counterfactual — research in parallel, debate findings, and synthesize a verdict in 22 seconds on AMD MI300X. The problem Understanding why a company failed usually takes weeks of research, and a single analyst's view is biased. What I built Forensic intelligence for failed companies. Twenty-four specialized agents work across four modes — Postmortem, Pre-Mortem, Founder Mode, Counterfactual — researching in parallel, debating their findings, then synthesizing one verdict. The whole parallel debate runs on a single AMD MI300X. The result A defensible verdict in 22 seconds, from 24 agents debating on one GPU. A showcase of multi-agent orchestration where agents argue toward a conclusion instead of a single model guessing. Multi-agent systems like this are the deep end of an AI build. Try other live agents in the AI Lab. ### Conductor https://anilpervaiz.com/work/conductor Cross-repo consistency agent. A plain-English change request becomes a fleet-wide migration: watsonx Orchestrate plans, IBM Bob Shell executes per-repo, an aggregator verifies fleet consistency. Real Bob sessions logged. The problem Rolling one change across dozens of repos is the kind of tedious, error-prone work that quietly eats senior-engineer time. What I built A cross-repo consistency agent. A plain-English change request becomes a fleet-wide migration: watsonx Orchestrate plans it, IBM Bob Shell executes per repo with framework-specific reasoning, and an aggregator verifies the whole fleet ended up consistent. Every Bob session is logged. The result One request migrated three frameworks (Express, Fastify, NestJS) with per-repo idiom reasoning, every session captured for review. A week of manual edits becomes a reviewed, auditable run. This is the kind of internal tool an AI build or retainer produces. More live demos in the AI Lab. ### DotChat https://anilpervaiz.com/work/dotchat PDF-to-cited-answer RAG product. Kimi K2.6 handles the chat, DeepSeek V4 Pro powers retrieval, Supabase pgvector stores the index. Compare mode runs both models in parallel on the same retrieved context. The problem Most "chat with your PDF" tools either hallucinate or hide which model and retrieval choices actually drive answer quality. What I built A PDF-to-cited-answer RAG product where every answer is grounded in specific pages. Kimi K2.6 runs the chat, DeepSeek V4 Pro powers retrieval, and Supabase pgvector stores the index — no Pinecone, no LangChain. A compare mode runs two models in parallel on the same retrieved context, so you can see exactly how model choice changes the answer. The result Every answer is page-grounded and citable, two models can be compared side by side, and the whole thing runs on pgvector instead of a managed vector vendor. A working demonstration that retrieval quality — not model hype — decides RAG quality. I wrote about this approach in building AI agents that don't hallucinate. Want a grounded knowledge agent? Book a call. ### SupportPilot https://anilpervaiz.com/work/supportpilot An embeddable AI support agent that answers from your help-center docs, escalates to humans, and writes ticket summaries. The problem Support teams drown in repetitive questions that are already answered in the docs — but customers won't read a help center, and hiring to keep up doesn't scale. What I built An embeddable AI support agent that answers strictly from your help-center content, escalates to a human the moment it's unsure, and writes a clean ticket summary for whoever picks it up. Retrieval is grounded so it can't invent policy, and the confidence threshold for handoff is tuned per client. It drops into any site with one snippet. The result In the first week it deflected 62% of tickets, with a median response under 8 seconds and a 4.7/5 satisfaction score — live one day after embedding. The team kept the hard conversations and handed off the repetitive ones. This is the AI Project Build tier in one artifact. The same retrieval-and-guardrails approach powers DotChat, and you can try a live agent in the AI Lab. ### Narrable https://anilpervaiz.com/work/narrable Brand and Webflow build for an AI-native voice notes app. Designed the marketing system, built the CMS, shipped in 18 days. The problem An AI-native voice-notes startup needed a marketing site that matched the polish of the product — and let a small team ship pages without a developer. What I built The brand and the Webflow build: a marketing design system, a CMS a non-technical team can run, and a component system that survives edits. Design to launch in 18 days. The result Trial sign-ups rose 318% in the first 30 days, the site scored 98 on mobile PageSpeed, and the team shipped on its own from day one. Craft and speed, not a trade-off. This is the craft foundation behind the AI work. Need a site like this? Book a call or see more work. ### Orbi RAG https://anilpervaiz.com/work/orbi-rag Internal knowledge agent with hybrid retrieval over Notion + Slack + Linear. Cuts onboarding research from days to minutes. The problem New hires lose days hunting for answers scattered across Notion, Slack, and Linear — and so does everyone they interrupt. What I built An internal knowledge agent with hybrid retrieval across Notion, Slack, and Linear. It grounds answers in the real source docs and is tuned for the messy, cross-tool reality of a working company's knowledge. The result 94% answer accuracy across 12,000 indexed docs, and onboarding research dropped from three days to half a day. Tribal knowledge became searchable. Hybrid retrieval like this is core AI architecture work. Want one for your team? Book a call. ### Pinokio Ops https://anilpervaiz.com/work/pinokio-ops n8n + Supabase automation chain that turns Stripe events into Slack alerts, finance exports, and CRM updates. The problem A growing team was losing hours every week to manual ops — copying Stripe events into Slack, spreadsheets, and the CRM by hand. What I built An n8n + Supabase automation chain that turns every Stripe event into the right Slack alert, finance export, and CRM update, automatically, with nothing falling through the cracks. The result 22 hours a week of manual ops removed, 9 workflows live, and zero missed events in 90 days. The team stopped being a human integration layer. This is the automation tier of an AI engagement. Want your stack wired together? Book a call. ### Saaspo Galleries https://anilpervaiz.com/work/saaspo-galleries Re-platformed one of the most-visited SaaS design galleries to Webflow CMS with a daily-update workflow. The problem One of the most-visited SaaS design galleries needed to move to a platform a small team could update daily, without engineering in the loop. What I built A full re-platform to Webflow CMS with a daily-update workflow — structured collections, constrained components, and an editing flow tuned for speed over flexibility. The result 1.2M monthly visitors served on a 100% Webflow-native build, with a 15-minute update workflow that keeps the gallery fresh every day. Volume at scale, run by a lean team. The discipline that keeps a 1.2M-visitor site fast is the same discipline I bring to AI builds. See more work. ### Aimiable Agent https://anilpervaiz.com/work/aimiable-agent Conversational AI for relationship coaching — orchestrated with Claude + tool use, packaged as a mobile-first PWA. The problem Relationship coaching is deeply personal — a conversational AI for it has to feel warm, stay cheap to run, and work on a phone. What I built A conversational coaching agent orchestrated with Claude and tool use, packaged as a mobile-first PWA so it installs to the home screen without an app store. The result 11,000 monthly active users, 92% day-7 retention on the paid tier, and a per-conversation cost of just $0.04 — proof that a thoughtful agent can be both warm and cheap to operate. Cost-aware agent design is part of every AI build. Try live agents in the AI Lab. ### OAG Brand System https://anilpervaiz.com/work/oag-brand Identity, type stack, and Webflow brand library for an aviation analytics company. From logomark to component kit. The problem An aviation analytics company needed one identity that worked across a marketing site and a data-dense product, without drifting between them. What I built A full identity system: logomark, type stack, and a Webflow brand library — 32 brand tokens that keep web and product visually in sync from a single source. The result One brand system across web and product, 32 tokens, delivered concept to handover in six weeks. Design tokens that behave like infrastructure, not a style guide nobody follows. Token-driven systems are how I keep both websites and AI builds consistent. See more work. ### Ritten CMS https://anilpervaiz.com/work/ritten-cms Marketing site + Webflow CMS for a clinical EHR product. Components designed for a non-technical content team. The problem A clinical EHR product needed a marketing site its non-technical team could fully own — in a regulated space where getting details wrong matters. What I built A marketing site and Webflow CMS built around a constrained component system: 12 reusable components with content-shape rules, so the team can't accidentally break the layout. The result 12 reusable components, 100% marketing self-service, and zero developer hours per month after launch. The team ships on its own, and the brand stays intact. The same systems discipline underpins the AI I build. See more work or book a call.