SaaS / 2026

A help-centre copilot inside a B2B SaaS product

An in-app assistant that answers customer questions from product documentation and past support tickets, with citations, permission-aware retrieval and a clean handoff to human support.

Sample project: an illustrative engagement showing how we scope and build this kind of work.

answers shown with source citations
100%
target time to first token
< 3s
click to hand off to a human agent
1

Context

This sample engagement describes a scheduling platform used by service businesses. Its help centre has several hundred articles, and its support team answers many tickets that the documentation already covers. Customers rarely search the help centre; they open a chat and wait for a person.

Challenge

The product team wanted an assistant inside the app that could answer “how do I” questions immediately, without inventing features that don’t exist. It had to respect plan tiers (a customer on the basic plan shouldn’t be told to use an enterprise feature), cite its sources, and pass the conversation to a human agent with full context when it couldn’t help. The in-house engineers were busy with the core roadmap and had no experience running LLM features in production.

What we built

We built the feature inside the existing React front end and Node.js API, working through pull requests reviewed by the product team:

  • Ingestion. A job that pulls help-centre articles and resolved support tickets, strips personal data from tickets, splits content by heading, and tags each chunk with product area and plan tier.
  • Retrieval. Hybrid search combining keyword and vector similarity, filtered by the customer’s plan before anything reaches the model.
  • Answering. A prompt that answers only from retrieved sources, cites each claim and says plainly when it doesn’t know.
  • Handoff. A button that opens a support ticket with the conversation, the sources used and the customer’s account context attached.
  • Feedback. Thumbs up and down on every answer, with optional comments, feeding a weekly review.

Architecture

Embeddings and chunks live in the product’s existing PostgreSQL database using pgvector, which avoided adding a new data store. Ingestion runs nightly and on article publish. The answering endpoint streams responses to the client and logs each request, retrieved chunks, model output and latency to Langfuse.

An evaluation set of real customer questions, each paired with the articles that should answer it, runs in CI. It checks retrieval recall, citation accuracy and refusals on out-of-scope questions. The model provider sits behind a small interface so the team can switch models without touching the feature code.

Stack

Claude API, PostgreSQL with pgvector, TypeScript, Node.js, React, Langfuse for tracing and evaluation, deployed on the client’s existing infrastructure behind a feature flag.

Have something you want built?

Tell us what you are trying to do and where it is stuck. We will reply with questions, an honest view on the right approach, and next steps.

Start a conversation