The constraint#
The UI is an afternoon. The requirement is not: the agent speaks as a representation of a real career, so a wrong answer can invent a job, a client or a skill that a recruiter believes.
The agent is intentionally constrained to the evidence in this portfolio and is tested against factual, attribution and safety regressions.
That is also why it does not roleplay as Brenda. It is her portfolio AI, it refers to her in the third person, and it says so when identity matters.
A truth layer before a personality#
Source-of-truth material is kept separate from prose. The agent distinguishes a fact it can support, a project currently being built, and something it simply does not know.
claim = retrieve(question) if claim.source_strength == "verified": answer_with_citation(claim) elif claim.source_strength == "in_progress": answer_with_status_label(claim) else: say_you_do_not_know()
Retrieval#
A recruiter question often crosses several documents, so retrieval runs over small, evidence-rich chunks carrying metadata: company, project, timeframe, technology, claim type. Postgres full-text search and pgvector embeddings are combined with reciprocal rank fusion, and a router step first decides whether a question needs evidence at all.
Retrieved text is treated as evidence, never as instructions. Instructions found inside a retrieved document are not followed.
A public agent is an adversarial surface#
Hard scope boundaries, with input and output checks.
Secrets kept entirely outside retrieval.
Server-issued sessions: three questions, and the count cannot be reset by refreshing.
Rate limits and logging for suspicious patterns.
Safe fallback behaviour.
The email collected before the demo is for access and abuse prevention, not a newsletter. Visitor identifiers are hashed server-side.
Evals#
Factual
Dates, titles, companies.
Boundary
Does not invent clients, revenue or scale.
Attribution
Keeps JPMorgan, Moody's and personal projects separate.
Status
Calls prototypes prototypes.
Retrieval
Cites the relevant case study, not the nearest keyword.
Language
EN and ES with the same facts.
The eval that matters most: does the agent make her sound more accomplished than the evidence supports?
Observability#
Each trace records the retrieved chunks, model latency, a token and cost estimate, safety decisions and groundedness. A bad answer stops being an anecdote and becomes a test case.
Architecture#
Browser widget ↓ streaming request Serverless API (Vercel) ├─ rate limit + server-issued session ├─ safety / scope check ├─ router: does this need evidence? ├─ hybrid retrieval │ ├─ Postgres full-text │ └─ pgvector embeddings │ └─ reciprocal rank fusion ├─ model generation (Claude) └─ trace + online scoring ↓ eval datasets / CI gate
Tracing and a prompt registry run on Langfuse, with a local fallback prompt so prompt retrieval is not a single point of failure. A GitHub Action pulls production failures into regression-test pull requests; a human merges, and prompts are never mutated autonomously in production.
What is running#
A static front end with a widget that holds no secrets.
Serverless API functions on Vercel for session creation and chat.
Hybrid retrieval over a Supabase Postgres corpus.
Online scoring of every answer for quality, groundedness and safety.
A CI-gated eval suite across six categories.
It was built with AI-assisted development tools, which is worth saying out loud on a page about not overstating things. Next: more evals drawn from real production questions, Turnstile, and voice once the facts hold.
Frequently asked questions#
Does the agent speak as Brenda?
No. It is Brenda's portfolio AI, refers to her in the third person, and says so whenever identity matters. A first-person persona makes a hallucination sound like a personal claim.
How does retrieval work?
Hybrid retrieval over a Supabase Postgres corpus: full-text search and pgvector embeddings combined with reciprocal rank fusion. A router step first decides whether a question needs portfolio evidence, and retrieved text is treated as evidence, never as instructions.
Why only three questions per visitor?
Abuse prevention on a public surface. Sessions are server-issued so the count cannot be reset by refreshing, and a server-generated visitor hash adds a second limit.
Was it built with AI assistance?
Yes. The code is hers and she can explain every part of it, but it was written with AI-assisted development tools rather than by hand alone.
Ask it something
The chat on this page is the system described above. It will tell you what it does not know.