Status: designed, not executed. These are original research scenarios, not claims about any competitor. They close the gap between documented design and observed app behavior. Run the same scenarios on Lose It!, MacroFactor Nutrition, Cronometer, MyFitnessPal and Cal AI; mark unsupported features N/A only after confirming the product's own statement or current UI.
Ground rules
Use a dedicated consenting test account, synthetic meal names and test data; never a person's real health history. Record app identifier, version/build, OS/device, country/storefront, account age, trial or paid tier, experiment cohort if visible, locale, timezone and capture date. Do not infer production behavior from old screenshots. Test Android first for Samarth's working context and iOS separately where platform differences matter. Do not initiate paid subscriptions without an explicit amount/term approval; cancellation and deletion require disposable accounts and separately confirmed scope.
For each run retain start state, a screen recording, ordered screenshots, taps, corrections, final stored values, reopen state, a result and unresolved observations. Record actual elapsed time from first entry action to correctly persisted result. Separate camera/network/model latency from interaction time; do not rank products using a single run or compare different devices without a caveat. Never label a photo estimate ground truth merely because it matches a nutrition label: portion identity and amount must also match.
Result vocabulary
- Pass: the defined research criterion was observed in the recorded version. This does not prove reliability across users.
- Fail: the observed result contradicted the criterion, with a reproducible path.
- Partial: part of the path works; name the missing state or unverified persistence.
- N/A — confirmed: explicitly unsupported or not applicable for this plan/platform.
- Blocked: access, permissions, account tier or environment prevented testing.
- Not run: no behavioral evidence yet. This is the default for this package.
A. First session and entitlement
T01 — Onboard without buying by accident
Start at clean install. Set a synthetic goal through ordinary onboarding. Record every requested input, why it appears to be needed, progress indicator, back/skip behavior and when the first useful screen becomes available. Inspect the trial/paywall rather than subscribing. Capture renewal cadence, total amount, dismissal, restore purchase and cancellation explanation. Pass criterion: the researcher can identify the chosen plan and next billing event without inference. If a subscription is essential, mark the subsequent test blocked pending authorization.
T02 — Goal revision and migration
Change the goal or target after setup. Record whether prior history remains intact, whether today/future/past targets change, and what explanation or confirmation appears. Undo if supported. Pass criterion: the user can distinguish target change from historical data mutation. Do not conflate a calorie recommendation with medical advice or algorithm quality.
T03 — Permission denial and recovery
Deny camera access, attempt scanning, then enable permission through the documented recovery. Repeat for notifications and health sync where applicable. Record fallback/manual path and whether entered work is preserved. Pass criterion: recoverable denial does not strand the user or silently erase their inputs.
B. Log one meal accurately
T04 — Database search, provenance and portion
Search the same generic ingredient and a packaged product. Record source badges, duplicate entries, nutrient completeness, default portion and gram entry. Log a weighed synthetic portion. Pass criterion: serving unit and amount are explicit, and the saved calories correspond to the selected record. Compare record quality separately from search convenience.
T05 — Barcode found / unknown
Scan a known item, then a genuinely unrecognized test barcode. Record camera behavior, rescan controls, which information carries into manual creation, and whether creating an item changes only the user's food or the shared database. Pass criterion: an unknown code has an explicit, recoverable route that does not misidentify a different item.
T06 — Label extraction with ambiguous serving
Use a test label containing both per-100g and per-serving information. Check calories/macros, decimal separators, serving count, unit conversion and editability before save. Pass criterion: the researcher can see and correct the basis of the extracted values. Record extraction output, not just the polished preview.
T07 — Photo estimate and transparent repair
Capture a controlled plate with known weighed ingredients, including an ingredient not visually obvious. Record detected food, quantities, uncertainty/provenance, and whether editing one component recalculates totals. Pass criterion: the proposed record is reviewable and editable; accuracy requires a separate evaluation set, not this single demonstration.
T08 — Voice/text multi-item interpretation
Use the same synthetic utterance, with a quantity correction in a second turn. Record whether it creates one meal or several foods, handles units, previews before saving, and mutates rather than duplicates the intended item. Pass criterion: the saved record matches the confirmed interpretation and cancellation leaves no ghost entry.
T09 — Correct an already logged entry
Change the portion, replace an ingredient, move meal/date, then delete and undo where supported. Close/reopen the app; check day totals, history and any exported record. Pass criterion: all exposed representations agree after the edit. Explicitly distinguish editing today's instance from editing a reusable template or global food record.
T10 — Repeated save / unstable connection
On a disposable account, test repeated taps on save and a single temporary network interruption. Record pending/success/error affordances and count persisted items after reconnection. Pass criterion: outcomes are understandable and duplicates can be detected/corrected. Do not perform denial-of-service or artificially high request-rate tests.
C. Returning-user workflows
T11 — Repeat yesterday's meal
Reuse a multi-item meal with one substitution. Compare recent/favorite/search/copy paths and date/meal defaults. Record whether changes affect the old day or the saved template. Pass criterion: the new instance is correct and original history remains as intended.
T12 — Batch recipe and cooked yield
Create a recipe, set total cooked weight and log a portion; then revise the recipe. Check whether historical portions retain original values or are recalculated, and whether the UI communicates this. Pass criterion: the relationship among ingredients, yield, serving and log instance is understandable. Do not assume all apps support cooked yield.
T13 — Recipe import correction
Import a public recipe permitted for testing; inspect ambiguous ingredients, optional ingredients, servings and substitutions. Record review/save friction and source link retention. Pass criterion: unresolved ingredient matches are visible before they affect the saved totals.
T14 — Weekly review and missing days
Use synthetic complete days plus a deliberately missing day. Inspect charts, averages and summaries. Pass criterion: zero intake, no data and partially logged days are distinguishable. For adaptive products, record how missing data affects eligibility and explanations; do not claim algorithm convergence from a short test.
T15 — Weight trend and target feedback
Add a noisy sequence of synthetic weigh-ins. Inspect raw and smoothed values, edit an older weigh-in, and review any proposed target adjustment. Pass criterion: the researcher can distinguish raw observations, estimated trend and recommendation, and identify user control over adopting a change.
D. Ecosystem and boundaries
T16 — Exercise and double counting
With a test health-data source, add an activity that may also arrive from a wearable. Record origin, import latency, deduplication, calorie adjustments and exclusion controls. Pass criterion: the user can identify whether exercise is counted once and how it changes the budget. Do not mix a separate workout app's features into the nutrition app's capability list.
T17 — Offline capability decomposition
Separately test opening cached history, searching a known food, searching an uncached food, saving a manual entry, barcode lookup and AI inference while offline. Reconnect and inspect sync. Pass criterion: the report can say which operations work, rather than using one misleading 'offline' checkbox.
T18 — Locale, accessibility and input ergonomics
Change serving-unit defaults and inspect decimal input, metric/imperial preferences, large text, screen reader labels, contrast, focus order and one-handed reach. Use OS accessibility tools rather than inferring accessibility from a screenshot. Pass criterion: primary logging and correction can be completed under the chosen accessibility condition; document the actual condition.
T19 — Export and portability
Export only synthetic test history. Compare entries, serving units, nutrient totals, timestamps, timezone, recipes, notes and edits against the app. Pass criterion: the format and lost information are documented. A CSV button does not prove useful portability or round-trip import support.
T20 — Cancel, restore and delete as separate flows
On an authorized disposable account, inspect cancellation of auto-renewal, restore purchase, sign-out, account deletion and data deletion separately. Capture retention/export warnings and whether store subscription management is separate. Do not assume deleting an app cancels a subscription. Pass criterion: the user can understand the consequences of each action; execute destructive steps only within the approved test scope.
Research log schema
run_id, app, bundle_id, app_version, OS, device, locale, timezone, storefront, account_tier, capture_date, scenario_id, start_state, evidence_paths, entry_actions, repair_actions, elapsed_seconds, saved_values, reopen_values, result, caveats, tester
Keep raw data private. Public reports should use synthetic examples, avoid account identifiers, and link evidence with capture date and platform. Every new run should update the atlas's specific unknown, not overwrite older evidence without a version note.