Workflow AppsWorkflow Apps

Case Study · Engineering

RAG Website Assistant for an Engineering Firm

A retrieval-augmented assistant for a national engineering firm's website: daily WordPress ingestion, vector search, cited answers, and edge deployment on Cloudflare.

RAG · Cloudflare Workers · WordPress · Claude · TypeScript

The Situation

A national environmental engineering firm runs a large WordPress site: hundreds of project, staff, news, location, and service pages. Visitors arrive with specific questions, like whether the firm has done a certain kind of remediation, which office covers a region, or who leads a practice area, and the answers are spread across pages that a menu doesn’t surface well.

The firm wanted an assistant on the site that answers those questions accurately, from its own published content, without inventing anything.

What We Built

A website assistant built on retrieval-augmented generation (RAG) and deployed as a Cloudflare Worker that serves both the chat widget and the API behind it.

  • Daily ingestion. A scheduled job pulls every public content type from the WordPress REST API (pages, posts, news, projects, events, locations, service areas, and staff), including structured custom fields. It converts each item to Markdown, splits it into heading-aware chunks, embeds them, and stores them in a vector index.
  • Retrieval before generation. Each question is embedded and matched against the index. Only passages above a relevance threshold reach the model, and the instructions require answers to come from that retrieved content alone.
  • Cited answers. Every response returns the source pages it drew from, and the widget links them under the answer.
  • Production guardrails. Origin allowlisting, per-visitor rate limiting, request validation, short-lived conversation memory, and an anonymized question log the marketing team can review.

Architecture Decisions

Edge-native, no servers to run. The Worker, vector index, embeddings, session store, and question log all run on Cloudflare. There is nothing to patch and no idle infrastructure to pay for.

Incremental, self-healing ingestion. Each run lists every page, so deletions and exclusions apply immediately, but only new, edited, or stale pages are re-embedded. A refresh budget keeps each run inside platform limits while still catching template and custom-field changes that don’t update a page’s modified date.

Model-agnostic by design. Requests go through an AI gateway with provider-specific formatting isolated in one place. The production model is Claude, and the team can compare models side by side on a test site before switching.

Fixing the source, not just the bot. Many of the site’s pages were rendered by page-builder templates, which left their REST content empty. We migrated 19 single-page templates onto their pages with a reversible WP-CLI script, added a must-use plugin to keep page-builder output from corrupting REST responses, and verified more than 400 project, staff, and news pages before and after the change.

What Changed

Visitors get direct, sourced answers instead of hunting through menus, and every answer can be checked against the page it came from. The content team keeps publishing in WordPress as before; the assistant picks up changes on its next daily run.

Have a Similar Project?

Tell us what you're working on. You'll hear back within one business day.