← All posts

Engineering · AI · Frontend

Shipping an LLM Feature in an Angular App

Mihajlo Petrović5 min read

Prototyping an AI feature takes an afternoon; shipping one takes longer, and almost none of the extra time is prompting. Streaming with SSE, cancellation, rate limits, caching and cost control - with the Angular and Node code.

Prototyping an AI feature takes an afternoon. Shipping one takes considerably longer, and almost none of the extra time goes into prompting — it goes into streaming, cancellation, error handling, and cost control.

This is the practical version for an Angular frontend with a Node backend: the architecture, the code that matters, and the four things that will bite you in production.


Rule Zero: The Key Never Touches the Browser

Every provider SDK will happily run in a browser and every one of them warns you not to. If the API key is in your frontend bundle, it's public — "it's minified" is not a security model. Someone finds it, and you're funding their side project.

So the shape is always:

Angular  →  your backend  →  the model provider

Your backend is where the key lives, and — more usefully — it's where you enforce auth, rate limits per user, the system prompt, input size limits, logging, and cost caps. None of that can live in a client you don't control.


Streaming, Because 20 Seconds of Spinner Is Unacceptable

A non-trivial response takes many seconds to generate. Waiting for all of it before showing anything makes a fast feature feel broken. Stream it.

Backend — the SDK gives you an async iterator; you forward text deltas to the client as Server-Sent Events:

import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();

app.post("/api/assist", async (req, res) => {
  res.setHeader("Content-Type", "text/event-stream");
  res.setHeader("Cache-Control", "no-cache");
  res.setHeader("Connection", "keep-alive");

  const stream = client.messages.stream({
    model: "claude-opus-5",
    max_tokens: 64000,
    system: SYSTEM_PROMPT,
    messages: [{ role: "user", content: req.body.question }],
  });

  req.on("close", () => stream.abort());

  for await (const event of stream) {
    if (event.type === "content_block_delta" && event.delta.type === "text_delta") {
      res.write(`data: ${JSON.stringify({ text: event.delta.text })}\n\n`);
    }
  }

  res.write("data: [DONE]\n\n");
  res.end();
});

That req.on("close") line is not optional. Without it, a user who navigates away leaves a generation running that you are still paying for.

Frontend — EventSource only does GET, so for a POST body use fetch with a reader:

@Injectable({ providedIn: "root" })
export class AssistService {
  private controller?: AbortController;

  async *ask(question: string): AsyncGenerator<string> {
    this.controller?.abort();          // one in flight at a time
    this.controller = new AbortController();

    const res = await fetch("/api/assist", {
      method: "POST",
      headers: { "Content-Type": "application/json" },
      body: JSON.stringify({ question }),
      signal: this.controller.signal,
    });

    if (!res.ok || !res.body) throw new Error(`Request failed: ${res.status}`);

    const reader = res.body.pipeThrough(new TextDecoderStream()).getReader();
    let buffer = "";

    while (true) {
      const { done, value } = await reader.read();
      if (done) break;
      buffer += value;

      const lines = buffer.split("\n\n");
      buffer = lines.pop() ?? "";
      for (const line of lines) {
        if (!line.startsWith("data: ")) continue;
        const payload = line.slice(6);
        if (payload === "[DONE]") return;
        yield JSON.parse(payload).text;
      }
    }
  }

  cancel() { this.controller?.abort(); }
}

Two details people skip: buffer across chunks (an SSE event can be split across two network reads — parsing per-chunk gives you random JSON.parse failures under load), and keep the AbortController so a new question cancels the previous one.

In the component, signals make the UI trivial:

answer = signal("");
streaming = signal(false);

async submit(question: string) {
  this.answer.set("");
  this.streaming.set(true);
  try {
    for await (const chunk of this.assist.ask(question)) {
      this.answer.update(a => a + chunk);
    }
  } catch (err) {
    if ((err as Error).name !== "AbortError") this.error.set("Something went wrong.");
  } finally {
    this.streaming.set(false);
  }
}

One caveat if you're on zone.js rather than zoneless: fetch streaming happens outside Angular's change detection in some setups. Signals handle this cleanly; if you're seeing text arrive but the view not update, that's your cause.


The Four Things That Bite in Production

1. Rate limits are a when, not an if

A 429 will happen. Handle it as a first-class path: exponential backoff with jitter, a couple of retries, and a plain-language message if it still fails. The SDK retries some failures for you — know which, so you don't stack your retries on top of its retries and turn one slow request into a thundering herd.

Also rate-limit your own users per account. One person holding down a button shouldn't be able to spend your monthly budget before lunch.

2. Cache the stable prefix

If your system prompt plus reference material is large and identical across requests, mark it cacheable. Repeat calls read the prefix at a fraction of the cost.

The catch: the prefix must be byte-identical. Inject new Date() into the system prompt and your cache hit rate is permanently zero, with no error to tell you. Check the cache-read token counts in the response to confirm it's working.

3. Costs are per-token and users are creative

Cap input length server-side. Cap max_tokens. Log token usage per request with the user id — the first time someone asks "why was the bill 4x last month", you want an answer, not a theory.

4. Failure is a product decision, not a catch block

What does the feature do when the model is unavailable? "Show an error" is one answer; degrading to a non-AI path is usually better. Decide deliberately — a feature that becomes a dead button during an outage is a feature that erodes trust in the whole product.


What I'd Do Differently Next Time

Build the eval set before the feature. I've now twice shipped something that worked on my five test questions and fell over on real ones — and without a fixed set of cases, every prompt tweak afterwards is a guess about whether things improved. That's a whole topic of its own, and I wrote it up separately.

The prompt is the easy part. The engineering around it is the feature.

  • #angular
  • #llm
  • #streaming
  • #typescript
  • #node
  • #ai

Written by

Mihajlo Petrović

Software engineer in Belgrade. Builds his own products and the AI automations that keep them running.

Have a task that repeats every week?

Tell me about it. If it can be automated well, I will show you how. If it cannot, I will say that too.

Tell me what to automate