NovaLifestyle Companion

Tests & validation

Everything a judge needs to verify Nova: recommender metrics against baselines, the organisers' scenarios run automatically against this live deployment, screenshots, and a two-minute self-test.

…live end-to-end checks passed
…judge scenarios passed (EN + AR)
16 / 16unit + starter tests passing
…recommender Precision@5 (holdout)

Test it yourself in 2 minutes

Open the app, enter a customer ID (e.g. U000002 or U024001), then click Test lab in the top bar. It holds every scenario from the brief as one-click steps, each listing what to check. Switch EN / عربي to test Arabic.

#Say or typeWhat you should seeCriterion
1(just log in)Nova opens proactively with a recommender pick and its reason (open cart, favourite, discount…)Proactive · Personalisation
2I'm preparing my new home. Can you help me choose what I need?Multi-category checklist (furnishings, kitchen, home care…) with total; asks for budgetUnderstanding
3I would like to buy a new smartphone. What would suit me?Only devices compatible with the customer's OS; personal reasons; product cardsRecommendation · Accuracy
4That's too expensive. My budget is now 100, and I don't want the first option.First option gone for good, all options ≤ 100, budget in Profile & memoryAdapt · Memory
5Add the second one to my basketCorrect item added; one complementary suggestion within budgetTools · Cross-sell
6Checkout → "Remove that item before placing the order."Summary card, then removal; the old summary is marked out of dateTools · Reliability
7Find something in this category within my budget.Same category, ≤ remembered budgetMemory
8Add it, checkout, then click Confirm order or say "yes, confirm" (or «نعم، أكد الطلب»)Order placed only now; reference shown; Orders tab updatedSimulated purchase
9Switch customer, then come backSeparate conversation/basket; "welcome back" recalls contextMemory · Separation
10Mic button or Voice mode (Chrome/Edge/Safari)Speech in EN or AR, spoken replies, same conversation as textVoice & text

Automated judge scenarios (live)

Loading…

Reproduce: python tests/e2e_scenarios.py https://companion.automagicdeveloper.com. Checks use structured facts (tool calls, basket, memory, orders, verified prices), not wording.

Recommender model validation

Holdout = 20% of train customers never seen by the model. Precision@5 averages over all customers; NDCG@5 uses binary gains. The companion uses out-of-fold predictions, so it never benefits from label leakage.

Precision@5 (holdout)

NDCG@5 (holdout)

Stability across 5 folds (P@5)

Candidate recall by source

Top features (share of gain)

Full report: artifacts/ml_report.md

Unit & starter tests

Companion (scripted fake LLM, deterministic)

  • ✅ confirm_order is never a model tool
  • ✅ relevant_items never reaches the companion
  • ✅ Search respects eligibility and budget
  • ✅ Rejection, budget change and "first/second option" references
  • ✅ Checkout requires the customer's own confirmation
  • ✅ Basket change invalidates the summary; negation blocks confirm
  • ✅ Customers are isolated

python -m pytest tests -q → 7 passed

Supplied starter package (unchanged)

  • ✅ 9 starter tests (basket, pricing, eligibility, checkout, idempotent confirm)
  • ✅ store.py, basket_tools.py, order_tools.py, tool_adapter.py used as supplied

python -m unittest -v → 9 passed

Screenshots

Captured automatically from the live deployment with a headless browser.

Nova · AI Lifestyle Companion · Gen AI Hackathon, AI Experts Conference 2026