Sample project: an illustrative engagement showing how we scope and build this kind of work.
- answers shown with source citations
- 100%
- target time to first token
- < 3s
- click to hand off to a human agent
- 1
Context
This sample engagement describes a scheduling platform used by service businesses. Its help centre has several hundred articles, and its support team answers many tickets that the documentation already covers. Customers rarely search the help centre; they open a chat and wait for a person.
Challenge
The product team wanted an assistant inside the app that could answer “how do I” questions immediately, without inventing features that don’t exist. It had to respect plan tiers (a customer on the basic plan shouldn’t be told to use an enterprise feature), cite its sources, and pass the conversation to a human agent with full context when it couldn’t help. The in-house engineers were busy with the core roadmap and had no experience running LLM features in production.
What we built
We built the feature inside the existing React front end and Node.js API, working through pull requests reviewed by the product team:
- Ingestion. A job that pulls help-centre articles and resolved support tickets, strips personal data from tickets, splits content by heading, and tags each chunk with product area and plan tier.
- Retrieval. Hybrid search combining keyword and vector similarity, filtered by the customer’s plan before anything reaches the model.
- Answering. A prompt that answers only from retrieved sources, cites each claim and says plainly when it doesn’t know.
- Handoff. A button that opens a support ticket with the conversation, the sources used and the customer’s account context attached.
- Feedback. Thumbs up and down on every answer, with optional comments, feeding a weekly review.
Architecture
Embeddings and chunks live in the product’s existing PostgreSQL database using pgvector, which avoided adding a new data store. Ingestion runs nightly and on article publish. The answering endpoint streams responses to the client and logs each request, retrieved chunks, model output and latency to Langfuse.
An evaluation set of real customer questions, each paired with the articles that should answer it, runs in CI. It checks retrieval recall, citation accuracy and refusals on out-of-scope questions. The model provider sits behind a small interface so the team can switch models without touching the feature code.
Stack
Claude API, PostgreSQL with pgvector, TypeScript, Node.js, React, Langfuse for tracing and evaluation, deployed on the client’s existing infrastructure behind a feature flag.
