Agent design
Deliverable 1: how the agent, tools, data sources and memory work together. Red marks the guarded purchase path.
Open full size: architecture.svg
What happens in one turn
Every message, typed, spoken or sent by a button, follows the same path, so voice and text always share context.
Text box, microphone (browser speech-to-text) or a UI button (Add, Not for me, Confirm).
If a checkout summary is pending and the customer's own words confirm it (EN or AR), the app places the order.
Profile, history digest, ML top picks, remembered preferences, shown options, basket, conversation summary.
The LLM calls tools (up to 8 rounds). Each step streams live to the UI ("Searching the catalogue…").
[[I0123]] tags become product cards with catalogue-verified prices. The shown options are remembered.Memory and transcript are saved per customer. Long chats are compacted into a running summary.
Tools
The six supplied starter functions run through tool_adapter.call_tool, with the customer ID bound from the session, never from model arguments. confirm_order is not visible to the model.
| Tool | Purpose | Notable behaviour |
|---|---|---|
get_recommendations | Recommender picks, filtered by category/price | Adds a "why" reason code; tops up with personalised search when filters are tight |
search_products | Catalogue search for the current need | Eligible only (region, OS, launch, owned); rejected items excluded; applies the remembered per-item budget; suggests an upgrade option |
get_product_details / compare_products | Exact facts and side-by-side comparison | Fit to budget and quality preference; cheapest / best quality / biggest discount |
get_complementary_products | Cross-sell | Mined co-acquisition complements plus category rules, capped by remaining budget |
plan_bundle | Multi-need plans ("new home", trip, fitness) | One pick per category, greedy down-grading until it fits the total budget |
update_customer_memory | Semantic memory writes | Goal, budget (per item or total), likes, rejections, style, quality, notes |
get_customer_insights | History deep-dive | Home/Shop/Rewards activity, open carts, favourites, returns |
get_basket, add_to_basket, update_basket, remove_from_basket, prepare_checkout, get_order | Supplied starter functions | Business errors are explained plainly. Any basket change invalidates a shown summary. add_to_basket returns cross-sell hints. |
confirm_order app-only | Create the simulated order | Called by the application after the Confirm button or an explicit customer confirmation. Idempotent. |
Memory & context
Working memory
Recent transcript window (text and voice turns, tool calls and results). Always cut at a user turn, so tool pairs are never split.
Episodic: sessions & summaries
Each login starts a clean conversation. The previous one is archived (folded and readable under "Previous conversations") with a recap of goal, products discussed, basket and orders. Nova uses these recaps to welcome the customer back and proactively pick up where they left off. Within a long conversation, older turns are folded into a running LLM summary.
Semantic preferences
Goal, budget and its scope, liked and rejected products, style and quality preferences, notes. Persists across logins, and "New chat" keeps it.
Referential
The ordered list of the last products shown, so "the first option" or "the second one" resolves deterministically.
Transactional
Pending checkout ID and orders. Cleared on any basket change, so a stale summary can never be confirmed.
Isolation & privacy
Everything is keyed by customer ID. Switching customer loads that customer's own conversation, memory and basket. "Forget my preferences" wipes semantic memory.
Models
Conversational agent: GPT (via LLM gateway)
gpt-6-lunathrough an OpenAI-compatible gateway with a dedicated app key, with automatic fallback togpt-5.6-terra.- Native tool calling; temperature 0.4 for conversation, 0 for summaries.
- GPT models only. No product fact is trusted from the model: names and prices come from the catalogue.
Recommender: LightGBM LambdaRank (ours)
- Candidates: own history, category affinity, declared interests, co-visitation, popularity trend, new arrivals (89% candidate recall).
- 83 features computed relative to each customer's snapshot date: time-decayed interactions, affinities, price vs budget, quality fit, discount, family and subscription crosses.
- Out-of-fold predictions for train customers (no label leakage). Holdout P@5 0.230 and NDCG@5 0.323, about 15× the baselines. Details →
Voice: browser-native
- Speech-to-text: Web Speech API (
en-US/ar-EG). Text-to-speech:speechSynthesis, with the voice picked by reply language. - Hands-free voice mode: listen → answer → speak → listen again. Voice turns get shorter replies.
- No audio leaves the browser for an LLM; the transcript joins the same conversation.
Arabic & English
- Full RTL interface with Arabic typography; the agent replies in the customer's language (MSA or Egyptian).
- Tool arguments stay in English enums, so Arabic requests search the catalogue correctly.
- Deterministic confirmation understands Arabic ("نعم، أكد الطلب") and blocks negations ("لا استنى").
Data pipeline
Built with DuckDB, streaming the CSVs out-of-core (1.09M interactions in about 7 seconds). relevant_items (ML labels) is dropped before anything reaches the companion.
| Source | Used for |
|---|---|
| train + test profiles (30,000 customers) | Personalisation: region and OS eligibility, household, membership, budget hint, quality preference, declared interests, owned items |
| products (1,200) | Facts, discounted prices, eligibility, comparisons, subscription flags |
| interactions (1.09M events) | History digest per customer: top categories, recent acquisitions, open carts, favourites not bought, returns, Home/Shop/Rewards activity. Also co-acquisition complements. |
Recommender output artifacts/recommendations.csv.gz | Top-20 per customer with reason codes; drives the opener, "For you" and get_recommendations |
Stack & deployment
Backend
Python 3.12, FastAPI, Server-Sent Events for live agent steps, httpx, SQLite for memory and the supplied simulation store (unchanged).
Frontend
Dependency-free single-page app (HTML/CSS/JS), Vodafone-style identity, EN/AR with RTL, Web Speech API, responsive down to phone width.
ML
DuckDB feature engineering, LightGBM LambdaRank, 5-fold out-of-fold scoring. Retrains end to end in about 4 minutes.
Infra
Docker Compose (companion-api + companion-web nginx) behind a TLS reverse proxy with Let's Encrypt. Portable to AWS (ECS/EC2) as is.
Quality
Unit tests with a scripted fake LLM (guardrails, memory, isolation), starter tests, and a live end-to-end judge-scenario suite. Results →
Source
github.com/m-abdelgawad/ai-lifestyle-companion, with a README covering run instructions and the demo script.