← All posts

Engineering · AI

AI Agents in Practice: BMAD, MCP, and SourceKit-LSP

Mihajlo Petrović9 min read

A practical look at the plumbing around AI coding agents: the BMAD method for planning, MCP for tooling, and SourceKit-LSP for real Swift code intelligence - what each one does, how to set it up, and where it actually pays off.

Most writing about AI coding agents stops at "it writes code for you." That's the least interesting part. The interesting part is the plumbing around the model: what it's allowed to read, what tools it can call, and what process it follows before it touches a file.

This post is the practical version. Concrete setups I actually run — the BMAD method for planning, MCP for tooling, SourceKit-LSP for real Swift code intelligence — and where each one earns its keep.


The Core Problem: Agents Are Blind by Default

A raw coding agent knows two things: your prompt, and whatever files it decided to open. That's it. It doesn't know your architecture, your conventions, or that the helper it's about to write already exists two folders over.

So it guesses. Confidently. And the output isn't usually broken — broken is easy to spot. It's plausible but wrong about your system: a swallowed exception, a duplicated service, a test that asserts the implementation instead of the behaviour. None of that fails a demo. All of it shows up eight months later.

Everything below is a different answer to the same question: how do I give the agent real context and a real process instead of hoping?


Layer 1: The Process — BMAD

BMAD (Breakthrough Method of Agile AI-Driven Development) is an open-source framework that stops treating the agent as one entity and splits it into role-based agents with distinct jobs — Analyst, PM, Architect, Product Owner, Scrum Master, Dev, QA. Each has its own persona, its own instructions, and its own output artifacts.

You install it into a repo:

npx bmad-method install

It drops agent definitions and templates into your project, wired for whichever IDE agent you use.

How the flow actually runs

Phase 1 — Planning (do this with a big-context model, before any code).

  1. The Analyst helps you produce a project brief: the problem, the users, the constraints.
  2. The PM turns that brief into a real docs/prd.md — functional requirements, non-functional requirements, epics broken into stories.
  3. The Architect reads the PRD and writes docs/architecture.md — stack, folder structure, data models, coding standards.

You review and edit both documents as documents. This is the whole trick: you're correcting a plan in plain English, where a wrong assumption costs one sentence to fix, instead of correcting a 400-line diff built on top of it.

Phase 2 — Development (in your IDE, one story at a time).

  1. The Product Owner shards the PRD and architecture into small, self-contained files.
  2. The Scrum Master drafts the next story — and critically, it embeds the relevant architecture context into the story file itself.
  3. The Dev agent implements only that story.
  4. The QA agent reviews against the story's acceptance criteria.

What actually makes it work

The mechanism isn't the personas — it's that each story file carries its own context. The dev agent isn't asked to infer your architecture from scratch every session; it's handed the slice that matters. That's what kills the "agent forgets the conventions after 20 minutes" problem.

When I use it — and when I don't

I use BMAD on greenfield features where the design isn't settled and the work spans more than a few files. The planning phase pays for itself there.

I do not use it for a bug fix, a dependency bump, or a two-file refactor. The ceremony costs more than the work. Running a full agile pipeline to rename a variable is a real failure mode, and it's easy to fall into because the process feels productive.

If you take one idea from BMAD without adopting the framework, take this: never ask for code before you've asked for a plan. "Read the deposit service and the existing cache usage. Propose two approaches with trade-offs. Don't write code yet." That one sentence has improved my output more than any prompt trick.


Layer 2: The Tools — MCP

MCP (Model Context Protocol) is the standard way to hand an agent tools beyond reading and writing files. An MCP server exposes capabilities — query this database, read this issue tracker, hit this API, run this linter — and any MCP-aware agent can call them.

Why it matters in practice: the difference between an agent that guesses your database schema and one that runs a query against it is the difference between a plausible migration and a correct one.

What's on in my setups, roughly in order of value:

  • The project's own docs and issue tracker. The agent reading the actual ticket beats me paraphrasing it.
  • Database/schema access, read-only. Real column names, real types, real nullability.
  • A browser. For agents doing frontend work, being able to load the page, read the console, and screenshot it closes the feedback loop. It stops being "this should work" and becomes "I checked."
  • Language servers. See below — this is the big one.

Two rules I hold to: read-only by default, and nothing production. An agent with write access to a live database is a story you tell at conferences, not a workflow.


Layer 3: Code Intelligence — SourceKit-LSP for Swift

This is the piece most people skip, and it's the one with the highest ceiling.

When an agent works on Swift, its default strategy is grep. It searches for a symbol name as text. That means it can't distinguish a definition from a mention in a comment, can't resolve an overload, can't follow a protocol conformance, and can't tell you what actually breaks if you change a signature.

SourceKit is the library behind Xcode's code intelligence — parsing, semantic analysis, refactoring, diagnostics. SourceKit-LSP wraps it in the Language Server Protocol, and it ships with the Swift toolchain:

xcrun sourcekit-lsp

It's what powers the VS Code Swift extension, and it gives you the queries a grep-driven agent is missing:

  • go-to-definition — the real one, not the first text match
  • find-references — every actual call site, type-aware
  • document/workspace symbols — the structure of a file without reading all 800 lines of it
  • diagnostics — compiler errors and warnings, live
  • hover — resolved types and signatures

The practical setup

Two things matter more than the wiring itself:

1. Build first, so the index exists. SourceKit-LSP relies on Swift's index-while-building store. On a project that's never been built, cross-file references are unreliable — for SwiftPM it's swift build, for an app target it's a real xcodebuild of the scheme. An agent asking for references against a stale index gets confidently wrong answers, which is worse than no answer.

2. Let it read structure instead of whole files. The genuine win isn't autocomplete — agents don't need that. It's that documentSymbol lets an agent understand a large file's shape for a fraction of the context that reading it costs. Context budget is the scarcest resource in any agent session, and structure-first reading is the cheapest way to stretch it.

For deterministic edits, the same ecosystem gives you better options than string replacement: SwiftSyntax for AST-level rewrites and swift-format to normalise the result. An agent that emits a syntax-tree transformation is doing surgery; an agent doing find-and-replace on Swift source is doing hope.

And the equivalent exists everywhere else — typescript-language-server, gopls, rust-analyzer. Swift is just where the gap between grep and semantics hurts most, because the type system carries so much of the meaning. It's the thing I most wish I'd wired up before my first serious SwiftUI project.


The Setup That Survives Contact With Real Work

Strip away the tooling names and this is what's left:

A project instructions file, checked into the repo. CLAUDE.md, .cursorrules, whatever your agent reads. Mine states the folder structure, the state-management approach, naming conventions, and — most valuably — the prohibitions: no new dependencies without asking, no any, no inline styles, no swallowed errors.

Point at real files, not at "best practices." "Follow the pattern in user-profile.service.ts" outperforms every abstract instruction I've tried.

One task, one branch, commit before and after. Reverting is then one command, and sunk cost is genuinely zero — which is the entire point of cheap generation.

A hard cap on diff size. If it's over a few hundred lines and I don't understand all of it, I throw it away and re-scope. Past that size my review degrades into skimming, and skimming is exactly how the subtly-wrong code lands.

Verification the agent performs itself. Run the tests. Load the page. Check the console. "It should work" is not a result.


Where This Actually Pays Off

Honest accounting of where the wins are — they're narrower than the marketing, but they're real:

  • Unfamiliar territory. An agent that explains idioms in terms of what you already know compresses the learning curve dramatically.
  • Mechanical refactors. Renaming across 40 files, migrating a deprecated API, converting components to standalone. Tedious, well-specified, easy to verify — and with a language server behind it, actually safe.
  • Reading legacy code. "Explain what this service does and where the side effects are" is worth an hour of my time on almost any day.
  • First-draft tests, especially the boring permutations I'd otherwise skip.

The pattern: agents are strongest where work is well-specified but tedious, and weakest where it depends on things about your system that were never written down. Every layer in this post is a way of writing those things down.


The Part That Doesn't Get Delegated

When something breaks at 9 a.m. and a few thousand people can't complete an onboarding flow, "the agent wrote it" is not an incident report. Whatever ships under your name is yours.

So the review bar goes up, not down — you didn't get the understanding that comes free with typing it yourself. My checklist hasn't changed in a year:

  1. Do I understand why every non-obvious line is there?
  2. Does this duplicate something that already exists?
  3. What happens on the failure path — really?
  4. Would I have written something structurally similar?

Number four catches the most. A "no" isn't automatically a rejection — sometimes the agent's approach is better than mine. But it always needs a reason.

BMAD gives the work a process. MCP gives it tools. SourceKit-LSP gives it truth about the code. Judgement is still the part you bring.


Running a different agent setup? I'm always curious what's working for other people — email me or reach out on LinkedIn.

  • #ai
  • #ai agents
  • #bmad
  • #mcp
  • #sourcekit
  • #swift
  • #developer productivity

Written by

Mihajlo Petrović

Software engineer in Belgrade. Builds his own products and the AI automations that keep them running.

Have a task that repeats every week?

Tell me about it. If it can be automated well, I will show you how. If it cannot, I will say that too.

Tell me what to automate