Test it yourself in 2 minutes
Open the app, enter a customer ID (e.g. U000002 or U024001), then click Test lab in the top bar. It holds every scenario from the brief as one-click steps, each listing what to check. Switch EN / عربي to test Arabic.
| # | Say or type | What you should see | Criterion |
|---|---|---|---|
| 1 | (just log in) | Nova opens proactively with a recommender pick and its reason (open cart, favourite, discount…) | Proactive · Personalisation |
| 2 | I'm preparing my new home. Can you help me choose what I need? | Multi-category checklist (furnishings, kitchen, home care…) with total; asks for budget | Understanding |
| 3 | I would like to buy a new smartphone. What would suit me? | Only devices compatible with the customer's OS; personal reasons; product cards | Recommendation · Accuracy |
| 4 | That's too expensive. My budget is now 100, and I don't want the first option. | First option gone for good, all options ≤ 100, budget in Profile & memory | Adapt · Memory |
| 5 | Add the second one to my basket | Correct item added; one complementary suggestion within budget | Tools · Cross-sell |
| 6 | Checkout → "Remove that item before placing the order." | Summary card, then removal; the old summary is marked out of date | Tools · Reliability |
| 7 | Find something in this category within my budget. | Same category, ≤ remembered budget | Memory |
| 8 | Add it, checkout, then click Confirm order or say "yes, confirm" (or «نعم، أكد الطلب») | Order placed only now; reference shown; Orders tab updated | Simulated purchase |
| 9 | Switch customer, then come back | Separate conversation/basket; "welcome back" recalls context | Memory · Separation |
| 10 | Mic button or Voice mode (Chrome/Edge/Safari) | Speech in EN or AR, spoken replies, same conversation as text | Voice & text |
Automated judge scenarios (live)
Loading…
Stability across repeated runs (current build)
Reproduce: python tests/e2e_scenarios.py https://companion.automagicdeveloper.com. Checks use structured facts (tool calls, basket, memory, orders, verified prices), not wording.
Recommender model validation
Holdout = 20% of train customers never seen by the model. Precision@5 averages over all customers; NDCG@5 uses binary gains. The companion uses out-of-fold predictions, so it never benefits from label leakage.
Precision@5 (holdout)
NDCG@5 (holdout)
Stability across 5 folds (P@5)
Candidate recall by source
Top features (share of gain)
Full report: artifacts/ml_report.md
Unit & starter tests
Companion (scripted fake LLM, deterministic)
- ✅
confirm_orderis never a model tool - ✅
relevant_itemsnever reaches the companion - ✅ Search respects eligibility and budget
- ✅ Rejection, budget change and "first/second option" references
- ✅ Checkout requires the customer's own confirmation
- ✅ Basket change invalidates the summary; negation blocks confirm
- ✅ Customers are isolated
python -m pytest tests -q → 7 passed
Supplied starter package (unchanged)
- ✅ 9 starter tests (basket, pricing, eligibility, checkout, idempotent confirm)
- ✅
store.py,basket_tools.py,order_tools.py,tool_adapter.pyused as supplied
python -m unittest -v → 9 passed
Screenshots
Captured automatically from the live deployment with a headless browser.