---
title: "How QueryLlama Works"
description: "From upload to answer: R2 storage, queued ingestion, sliding-window chunking, Workers AI embeddings, pgvector search and the planned citation layer."
url: "https://royalsoftworks.com/products/queryllama/how-it-works/"
source: "https://royalsoftworks.com"
format: "markdown"
note: "Markdown rendering of the HTML page at `url`. Same content, same canonical URL."
---

How it works

# From upload to answer, step by step.

A document doesn't sit passively in storage — it goes through a pipeline that makes it searchable within moments of arriving.

The pipeline

## What happens when a document is uploaded.

1. 01

   ### Upload goes straight to object storage

   The client uploads directly to a Cloudflare R2 bucket via a presigned URL — the file never transits the API server itself.
2. 02

   ### A job is queued

   The API enqueues a processing job on Cloudflare Queues, and an ingestion worker picks it up.
3. 03

   ### The document is split into overlapping chunks

   Text is divided into a sliding window of roughly 2,048 characters per chunk, with a 400-character overlap so a sentence split across two chunks doesn't lose context.
4. 04

   ### Each chunk is embedded

   Cloudflare Workers AI generates a 768-dimension vector embedding per chunk, capturing its meaning, not just its words.
5. 05

   ### Everything lands in Postgres

   Chunks, their embeddings (via pgvector), and a generated full-text-search column all land in the organization's scoped rows, ready to be queried the moment indexing finishes.
6. 06

   ### A query searches by keyword, meaning, or both

   Keyword search hits the full-text index; semantic search compares your question's embedding against chunk embeddings; hybrid search blends both rankings with reciprocal rank fusion.
7. 07

   ### (Planned) An answer is generated with citations

   The retrieval-augmented generation layer will take the top-ranked chunks, generate an answer, and attach a citation back to the source file and passage — this step is still being built.

For your IT team

- **Storage** — Cloudflare R2, presigned direct upload
- **Queue** — Cloudflare Queues for ingestion jobs
- **Chunking** — ~2,048 characters per chunk, 400-character overlap
- **Embeddings** — 768-dim vectors, Cloudflare Workers AI (BGE base), stored in pgvector
- **Database** — Supabase Postgres; Row-Level Security scopes every query to the caller's organization

02 Architecture

## A hosted pipeline, not a desktop app.

A document you upload goes to object storage, gets picked up by an ingestion worker that splits it into overlapping chunks and embeds them, and lands in Postgres ready to be searched — by keyword, by meaning, or both.

For your IT team

- **Chunking** — Sliding window, ~2,048 characters per chunk with 400-character overlap
- **Embeddings** — 768-dimension vectors via Cloudflare Workers AI (BGE base), stored in pgvector
- **Keyword search** — Postgres full-text search over a generated tsvector column

## Want the full technical reference?

Data model, RLS policy structure, and a capability-by-capability build-status breakdown — for your security and procurement review.

[Read the whitepaper](https://royalsoftworks.com/products/queryllama/whitepaper/) [Security](https://royalsoftworks.com/products/queryllama/security/)
