Home
Evidence base
Each entry below names a feature, the research under it, and a link you can check. Where the evidence runs out, or is not about our situation, there is a separate mark — not a footnote, part of the entry.
22 entries, 17 of them carrying a “this is where the data ends” mark, 74 linked sources.
The same text a person sees in the app, under “Evidence”. This page and that screen read one file, so they cannot drift apart.
🧷 If-then reminders
Strength meta-analysis
What the feature does
A reminder anchored to another event ending, delivered as the cue sentence itself: “Lunch just ended — now: call the clinic.”
What it rests on
This is the implementation-intention format. Across prospective-memory studies it improves follow-through with d = 0.508 (95% CI 0.419–0.597); a meta-analysis of 642 tests reports d = 0.36 overall — d = 0.27 for behaviour change, and d = 0.15 once publication bias is corrected — and finds effects are larger when the plan keeps its if-then shape (d = 0.43, against d = 0.29 for a plain schedule) and is rehearsed at least once — sending the cue sentence at the cue moment is that rehearsal. Replicated in ADHD specifically: an if-then plan lifted response inhibition in children with ADHD to the level of children without it.
Sources
💜 Emotion wheel diary
Strength meta-analysis
What the feature does
Pick a feeling from a granular wheel, answer “I don’t know” with no penalty, and mark where in the body it sits.
What it rests on
Naming a feeling is itself regulation: affect labeling produces the experiential, autonomic and neural signature of deliberate emotion regulation (lower amygdala response, higher prefrontal activity) while feeling effortless — and choosing among offered labels counts. Finer granularity predicts less maladaptive coping and less severe anxiety and depression. The penalty-free “I don’t know” exists because alexithymia affects roughly 50% of autistic people versus about 5% of non-autistic controls: “I can’t name it” is the normal state for much of this audience, not a failure to comply. Body-area marking rests on bodily emotion maps (n = 701, five experiments) that are statistically separable and consistent across cultures.
Where the evidence stops
Interventions aimed at alexithymia move “describing feelings” (g = −0.39) and “externally-oriented thinking” (g = −0.45) — and do NOT move “identifying feelings” (g = −0.06, p = .43, k = 16, N = 1215, one-tailed tests). No study has tested an emotion wheel as an interface. A body map narrows the field of candidates (72% in a one-emotion-against-all discrimination) rather than naming the emotion (full classification 38%, fear at chance). So this is a tool for putting a word to something already noticed — not for working out what is going on.
Sources
🏁 Duration calibration & finish line
Strength meta-analysis
What the feature does
The app records planned-vs-actual duration, then offers your own correction factor when you schedule and projects when a plan really finishes.
What it rests on
The planning fallacy — underestimating your own tasks even while knowing past ones ran long — is corrected most reliably by the “outside view”: using a reference class of comparable finished cases instead of the details of this one. A per-category actual/planned ratio built from your own completed events is exactly that personal reference class. A second, independent lever is unpacking: listing a task’s substeps before estimating measurably reduces underestimation. ADHD stacks a measurable timing deficit on top of the general bias.
Where the evidence stops
An analysis of 47 popular task apps (2024) found these strategies are almost never implemented in shipping products — so this is well-evidenced psychology, but rarely-tested product design.
Sources
🧠 Voice brain dump
Strength meta-analysis
What the feature does
Speak everything on your mind; it gets split into separate tasks with proposed free slots, so capture costs one tap and no typing.
What it rests on
Working-memory impairment is a meta-analytically established feature of ADHD (pooled effects around d = 0.69–0.74), and the standard clinical model prescribes moving information out of working memory into an external store — which is what capture-first design does. Offloading’s advantage over internal memory grows with load. Voice as the channel is a measured choice: speech input was 3.0× faster than a state-of-the-art phone keyboard in English (2.8× in Mandarin) with 20.4% lower error rate.
Sources
🌱 Lifts & drains
Strength meta-analysis
What the feature does
A quick state check right after an event; over time it shows which recurring commitments lift you and which drain you.
What it rests on
This is activity scheduling / behavioural activation in calendar form — monitoring mood alongside activities, then shifting time toward what pays back. A meta-analysis of 16 studies (780 participants) found a large pooled effect versus control (d = 0.87, 95% CI 0.60–1.15), and no significant difference from cognitive therapy across ten head-to-head studies — meaning the monitoring-and-scheduling component carries the effect on its own.
Where the evidence stops
That evidence comes from clinical depression samples, not general or neurodivergent productivity users. Our card is a self-insight surface, not delivered therapy.
Sources
🌅 Morning day-check
Strength meta-analysis
What the feature does
One morning push only when the day actually contains an overlap or an overload — silence otherwise.
What it rests on
The planning half rests on implementation intentions: a meta-analysis of 94 independent tests found a medium-to-large effect on goal attainment (d = 0.65) from settling when, where and how in advance. The quiet-unless-there-is-a-problem rule is the countermeasure to alert fatigue — repeated low-value alerts train people to dismiss everything, with override rates of 49–96% in clinical systems, so a push that fires regardless of content would destroy the one that matters.
Sources
🍎 Body fuel
Strength meta-analysis
What the feature does
Spots a day with no midday gap or a 4h+ unbroken booked stretch and tells you to protect a break. It guards TIME only — it never asks about food.
What it rests on
A systematic review and meta-analysis of 22 studies found micro-breaks reliably reduce fatigue and increase vigour (performance gains needed longer breaks). Why it matters more here: autistic burnout is characterised across 48 studies (~4,000 participants) as debilitating exhaustion driven by demand without recovery, and roughly half of autistic people meet the cut-off for alexithymia — 49.9% against 4.9% of non-autistic comparisons — so the internal “I need a break” cue is a less reliable thing to wait for. Interoceptive accuracy is not the reason: a 2025 meta-analysis of five adult studies found no significant autistic/non-autistic difference in it (p = .06).
Where the evidence stops
The refusal to track food is a deliberate design line: calorie- and fitness-tracking use is linked to disordered-eating symptoms (correlational, mostly non-autistic samples).
Sources
⏳ Living dial
Strength controlled trial
What the feature does
The event you are in drains in real time — the spent part dims — and Focus mode shows what is left as a shrinking wedge.
What it rests on
Impaired time perception in ADHD is measured, not a habit — but it is an average gap between groups, not a per-person trait: a 2024 meta-analysis pooling 824 effect sizes across the lifespan found a mean Hedges g ≈ 0.688, moderated by the age of the sample. A 55-study meta-analysis found deficits in all four standard paradigms and a different one in each: discrimination worst for sub-second durations; estimation and production accuracy consistent with an accelerated internal clock; and in reproduction, slower counting at short intervals (distraction) and faster counting at long ones (delay aversion). The adult evidence is the thin part — a review of the decade to 2022 found only 9 adult studies, with mixed results. On making the remainder visible: 44 French children aged 7–9 sat a 5-minute timed maths test twice, once with a visible Time Timer disc and once without. Anticipatory anxiety, measured after the instructions and before the task, was lower with the timer (t(43) = 2.77, p = 0.008, d = 0.42), as was inattentive behaviour during it (p = 0.002, r = 0.42); an earlier study from the same lab, in 49 children, found a visible timer beat a hidden one on the same measure (d = 0.40). Task accuracy did not change (p = 0.66): the gain is regulation, not speed. The extra benefit for children scoring higher on an ADHD screening scale held for inattention only (ρ = 0.448, p = 0.013, among the 30 children with screening data) — not for anxiety, where the correlation was −0.21 and not significant.
Where the evidence stops
The pooled g ≈ 0.688 is an average gap between groups whose distributions overlap heavily, so plenty of people with ADHD judge durations normally; the effect is moderated by age, and the adult base is thin — 9 studies in the decade to 2022. The timer evidence sits closer to our format than most on this screen: a Time Timer is a disc that drains as time passes, like Focus mode’s wedge. But it counted down a single 5-minute deadline set by an adult, for 7–9-year-olds in France — not a day-long dial, and not a deadline you set yourself. “Higher ADHD risk” there is a score on an abbreviated 10-item Conners observer scale, not a diagnosis, and that finding is a correlation among 30 of the 44 children rather than a randomised comparison.
Sources
🕐 The Dayring
Strength controlled trial
What the feature does
The whole day as a circular dial, each event a proportional arc — so “how much is left” is a shape, not arithmetic.
What it rests on
Randomized trials show that training time-processing skills together with external time-visualising devices improves time-processing ability and everyday time management — in children with ADHD (n = 38) and in a cluster-randomised trial with children with intellectual disability. Visual activity schedules are separately meta-analysis-backed for independent task sequencing.
Where the evidence stops
No trial has tested a circular 24-hour dial specifically. The devices in those trials were linear or analog time aids — the ring format is our extrapolation from an evidenced principle.
Sources
🔕 Calm notifications
Strength controlled trial
What the feature does
Every push type has its own switch, plus quiet hours — so notification volume is chosen, not inherited.
What it rests on
In a randomized field experiment (n = 237), batching notifications into predictable intervals left people more attentive and productive, in better mood and less stressed than as-usual delivery. Notably, the zero-notification group reported HIGHER anxiety and fear of missing out — which is the argument for a per-type matrix instead of one global mute. The reason volume must be controlled at all is alert fatigue: in clinical systems 49–96% of interruptive alerts get overridden as people desensitise.
Sources
✍️ Your word first, suggestions after
Strength controlled trial
What the feature does
A rule the check-in is being built to, named here before it holds everywhere. Your own attempt comes first: no suggested label, from the app or from AI, appears above it. Suggestions never arrive alone — at least three at a time plus an explicit “none of these” — and never with a model confidence score.
What it rests on
A suggested label shifts what you report, and you cannot feel it happening. In a randomized experiment (n = 1401), after interacting with a biased AI model the share of one particular answer in an emotion-aggregation task rose from 49.9% (±1.1) to 56.3% (±1.1), p < .001, d = 0.84; across blocks 50.72% → 61.44%, b = 0.02, t(50) = 6.23, p < .001. The same interaction with HUMANS moved it from 50.6% to 51.45%, p = .48, d = 0.10 — effectively nothing. The AI shifted people harder than people did because it was an amplified version of their own bias (human raters 53.08%, the CNN trained on them 65.33%). Critically, participants did not notice: their subjective estimate of being influenced did not match the objective shift, p = .90, d = −0.01. Two independent lines point the same way: fixed-option (forced-choice) formats systematically inflate apparent emotion “recognition” compared with free verbal labeling, and after semantic satiation of an emotion word people categorise emotional faces more slowly and less accurately — language does not just describe perception, it takes part in constituting it.
Where the evidence stops
The experiment ran on aggregating other people’s faces, not on your own state, and not in an alexithymic sample. The direction of the risk is confirmed by two independent lines; the size of it for our format is unknown. So the ordering rule is a precaution we chose to pay for, not a measured effect of this interface.
Sources
🎚️ When naming does not help
Strength controlled trial
What the feature does
Planned, and not in the check-in yet: intensity will be asked first and act as a gate — when the signal is faint the wheel will not open by itself, an action is offered instead, and “pick a word anyway” stays one tap away. Today the wheel opens straight up, for any intensity.
What it rests on
Naming a feeling does not help unconditionally — it depends on how strong the feeling is. In a randomized lab experiment, Study 1 (N = 63) found a main effect of strategy, F(1,60) = 13.20, p < .001, η²p = .18, and no effect of WHEN the label was applied, F(2,60) = .26, p = .97. Study 2 (N = 79) found a strategy × intensity interaction, F(1,76) = 22.37, p < .001, η²p = .23: at low intensity (arousal M = 5.02) labelling the emotion INCREASED distress, t(78) = −2.6, p = .011, while at high intensity (arousal M = 6.15) it lowered distress, t(78) = 2.4, p = .018. The proposed mechanism is semantic priming activating the negative emotional network, so on a faint state the label amplifies the negative rather than damping it. This is a property of the label, not of the person using it.
Where the evidence stops
Participants chose between only TWO labels, the stimuli were IAPS pictures, and the sample was psychiatrically healthy — not alexithymic, not autistic, not traumatised. No replication found. Carrying that over to a wheel of 49–74 labels for your own state is our caution, not a proven rule for this format — which is why the gate never blocks anything: “pick a word anyway” is always one tap away.
Sources
🔤 Readability: spacing, not fonts
Strength controlled trial
What the feature does
A two-step setting that increases line height, letter spacing and text size across most of the app — spacing, deliberately not a special “dyslexia font” — plus a button to have the day’s events read aloud. A dyslexia font (OpenDyslexic) is available as a separate opt-in, labelled in Settings as a personal preference rather than an accessibility feature, for the reason below.
What it rests on
In a crossover experiment, 74 dyslexic children (34 Italian, 40 French, ages 8–14) read the same text twice two weeks apart — once normally spaced, once with letter spacing nearly doubled plus matching increases in word and line spacing. The spaced version halved reading errors and raised reading speed (accuracy F(1,70) = 35.16, speed F(1,70) = 27.96, both p < .0001); separating spacing from the practice effect of re-reading leaves about 0.3 syllables/second, which the authors equate to the average improvement across one school year for Italian dyslexic children. In the paper’s more conservative test — 30 of those Italian dyslexics (mean age 10.9) against younger normal readers matched for reading level and IQ (mean age 7.8) — spacing moved the dyslexic group from 1.75 to 1.94 syllables/second (p = .001) while the matched controls did not reliably shift (1.87 vs 1.93, p > .3); the dyslexia-specific advantage there was significant for errors (group × spacing F(1,58) = 5.95, p = .018) and only marginal for speed (F(1,58) = 2.81, p = .09). The authors attribute the benefit to dyslexic readers being abnormally affected by visual crowding between letters. The other half of “not fonts” is tested just as directly: two groups of 64 children each, with and without dyslexia, read 8 equivalent texts, and “the data collected failed to show any effect from the letterform”. A 2026 meta-analysis pooling 15 studies of dyslexic children (91 effect sizes, N = 688) puts “dyslexia-friendly” fonts at g = −0.04, 95% CI [−0.15, 0.07], p = .5 — indistinguishable from zero. Preference does not rescue them either: 170 dyslexic children chose Arial over Dyslexie (44.7% vs 37.6%, and 61.9% vs 20.2% in the reversed order), and among 102 dyslexic plus 45 typical readers only 11.6% preferred Dyslexie, against 38.1% for Arial and 29.9% for Times New Roman (χ²(3) = 23.92, p < .0001). Reading the day aloud rests on a meta-analysis of text-to-speech and read-aloud tools for students with reading disabilities, showing a real comprehension benefit (d = 0.35, 95% CI 0.14–0.56, p < .01).
Where the evidence stops
The tested spacing nearly doubled letter spacing AND increased word and line spacing to match, in print, for children already diagnosed with dyslexia. Our toggle raises letter spacing by a small fixed amount and line height by 1.5×, but not word spacing — the mobile platform has no such property, so we did not fake it. A follow-up study found that widening letter spacing WITHOUT a matching word-spacing increase can slow reading rather than help it, in both dyslexic and typical readers. Our specific numbers sit closer to that untested combination than to either study’s tested one — a reasonable bet, not a validated one. Whether the spacing gain is specific to dyslexia is contested in print: a PNAS Comment argued the control group’s null could be a power artefact, and the authors replied. And the font evidence, strong as it is, is all children reading passages and word lists — not adults reading a calendar. It is enough to keep the font out of our accessibility answer; it is not enough to say it has been ruled out for you.
Sources
⏭ Transition alerts
Strength clinical guideline
What the feature does
About 5 minutes before an event ends, if the next starts within 15 minutes, a push names what to wrap up and what is next.
What it rests on
Advance, concrete transition warnings are among the longest-standing supported practices in autism. Visual transition supports significantly reduced the latency between the instruction and actually starting the next activity in a reversal-design study. The national evidence review that screened 31,779 abstracts and 972 studies lists both antecedent-based intervention and visual supports among its 28 evidence-based practices — the category this alert belongs to.
Where the evidence stops
The primary trials are in children, not adults, and the specific “wrap up X → next Y” wording is practice consensus carried over from first-then boards, not an independently tested component.
Sources
💊 Medication reminders
Strength clinical guideline
What the feature does
Escalating reminders for a dose, and marking it taken late still counts rather than being recorded as missed.
What it rests on
The NICE guideline on ADHD explicitly recommends visual reminders — apps, alarms, clocks, pill dispensers, calendar notes — to support regular medication taking, which puts app reminders inside the clinical guideline rather than only in app-design literature. A 2024 mixed-methods systematic review with meta-analysis found a significant improvement in adherence (OR = 2.39, 95% CI 1.19–4.79).
Where the evidence stops
Not uniform: a separate pooled analysis of four trials found a non-significant effect (OR = 2.32, 95% CI 0.91–5.90). The honest read is “probably helps, wide confidence intervals”. That a late dose still counts is our design decision, with no direct trial behind it.
Sources
🥄 Energy budget & pacing
Strength clinical guideline
What the feature does
Give any event or category a 0–5 energy cost (“spoons”), set a daily budget, watch it deplete as a ring in the dial’s centre, and get flagged in the morning check when today’s plan spends more than you have.
What it rests on
NICE’s ME/CFS guideline (NG206, 2021) recommends energy management to everyone with ME/CFS — its own term for “a self-management strategy that involves a person with ME/CFS managing their activities to stay within their energy limit” — where an exercise programme is gated: recommendation 1.11.10 says to only consider one, and only for people ready to progress beyond their activities of daily living or who want to add exercise. Recommendation 1.11.2: discuss the principles, explaining that energy management “is not curative”, “includes all types of activity (cognitive, physical, emotional and social)”, and “uses a flexible, tailored approach so that activity is never automatically increased but is maintained or adjusted”. Recommendation 1.11.3: help the person build a plan that records seven kinds of load — cognitive activity; mobility and other physical activity; ability to undertake activities of daily living; psychological, emotional and social demands; rest and relaxation, both quality and duration; sleep quality and duration; and environmental factors including sensory stimulation. Recommendation 1.11.7: “make self-monitoring of activity as easy as possible by taking advantage of any tools the person already uses, such as an activity tracker, phone heart-rate monitor or diary.” That is a real clinical guideline behind budgeting load before you overspend it, not a wellness-app invention. It matters beyond ME/CFS too: a systematic review of 48 studies (30 qualitative, 7 quantitative, 11 mixed methods; about 4,000 autistic people) describes autistic burnout as debilitating exhaustion and increased disability, chronic with intermittent crises, driven by sensory and social overwhelm, camouflaging, stigma and everyday demands. Reading your own internal state is measurably harder: a meta-analysis of Toronto Alexithymia Scale studies found, across the 9 reporting caseness, 49.9% of autistic participants above the alexithymia cut-off against 4.9% of non-autistic ones (risk ratio 6.50, 95% CI 3.26–12.93), and across 11 studies (292 vs 275 participants) difficulty identifying feelings higher by d = 1.28 (95% CI 0.96–1.60). That is the case for making the budget external and visible rather than trusting it to be felt. Interoceptive accuracy is NOT part of that case and we will not claim it is: a 2025 meta-analysis of five adult heartbeat-perception studies found no significant difference (−0.21, SE 0.11, 95% CI −0.43 to 0.01, p = .06), its authors concluding that autism is not systematically associated with altered cardiac interoceptive accuracy. The “spoons” language itself traces to Christine Miserandino’s 2003 essay on chronic-illness energy limits and has since become common shorthand in ADHD/autism communities for a fixed daily capacity that activities draw down.
Where the evidence stops
NICE evidences energy management as a strategy, not any numeric scale or ring visualisation — no trial has tested a 0–5 “spoons” rating or a depleting dial centre, and spoon theory itself is a patient-coined metaphor, never an instrument validated against objective fatigue or output. Both the scale and the ring are our design, built on an evidenced principle. We also budget less than NG206 plans for: our cost sits on scheduled events and categories, so of its seven load types, rest and relaxation, sleep, and environmental and sensory load never enter the number — the app reads sleep elsewhere, but not into this budget. The burnout review is a thematic synthesis of mostly qualitative work, and its samples were predominantly White, female, late-diagnosed adults with at least average intellectual or verbal ability. Alexithymia is measured for emotions, not for energy: we treat it as a reason not to rely on a felt signal, not as evidence that autistic people cannot feel tiredness.
Sources
🌫️ Two kinds of “I don’t know”
Strength observational
What the feature does
The store already keeps the two apart; the check-in is being split to match. “Something’s there, I can’t name it” and “Nothing right now” become two separate answers instead of today’s single “I don’t know”. Both are complete entries — neither is a skip, and neither breaks the streak.
What it rests on
They are different data. In a week-long ecological momentary assessment study — 29 autistic adults without intellectual disability against 28 non-autistic adults matched on age, sex and education, labelling emotions by phone — the autistic group chose “I have an emotion I can’t name” significantly more often, and reported mixed (simultaneously positive and negative) emotions more often. In both groups, absence of labelling and intense negative emotion went with poorer emotion control; but the link between “no emotion” and poorer control existed ONLY in the autistic group. Collapsing both answers into one “unsure” flag would erase exactly that distinction.
Where the evidence stops
The study shows that people USE these options and that they behave differently — not that having the option improves anything. Observational EMA, small groups, an autistic sample not selected for alexithymia (the two overlap but are not the same thing), and no causal claim. The full text is paywalled, so the numbers come from the abstract and exact coefficients are not available.
Sources
🔥 Load heatmap
Strength observational
What the feature does
Tints each day in a week/month grid by how heavily booked it already is, so accumulating load is visible before you live it.
What it rests on
The deficit it addresses is well evidenced: 55 studies find consistent time-perception impairment in ADHD, and in adults time reproduction runs shorter than reality — people systematically under-feel how much a day already holds. On the autism side, a 48-study review characterises burnout as demands accumulating over time with insufficient recovery, improving when demands are cut before crisis.
Where the evidence stops
Say it plainly: no study tests a day-load heatmap. The deficit is evidenced; this particular visual format is untested design.
Sources
✅ Forgiving streaks
Strength observational
What the feature does
A completion streak that absorbs a missed day and never resets to zero or frames the gap as a loss.
What it rests on
A 12-week real-world habit study (n = 96) concluded that missing a single opportunity did not materially affect habit formation — so resetting a streak to zero punishes something the behavioural data says is harmless. The harm model is the abstinence violation effect: attributing a lapse to personal failure (guilt, shame, hopelessness) is what turns a lapse into a relapse. Compassionate framing is supported by a meta-analysis of 27 randomized trials of self-compassion interventions.
Where the evidence stops
No trial tests grace-shield streak mechanics in an app — that specific design is our extrapolation.
Sources
🚩 Moment flag
Strength observational
What the feature does
One tap timestamps that something happened, with no fields to fill; you name and make sense of it later, when you have capacity.
What it rests on
This applies the ecological-momentary-assessment principle to self-tracking: EMA exists because retrospective self-report is distorted by recall bias — people systematically over-estimate both positive and negative affect in hindsight — and capturing in the moment minimises that distortion. Deferring the naming step is cognitive offloading, whose advantage over internal memory grows with load.
Where the evidence stops
EMA’s advantage is a measurement finding — more accurate data — not a demonstrated therapeutic effect of a one-tap flag.
Sources
🔁 Body-signal echo
Strength practice consensus
What the feature does
When a body signal recurs, it shows you the words YOU used for a similar signal before — instead of generating a fresh interpretation.
What it rests on
This is the weakest-evidenced feature we ship, and we would rather say so than dress it up. It is the software form of an interoception curriculum, where a person builds their own dictionary mapping body sensations to emotion names. Support is small-scale only: an 8-week pilot with eight autistic students, and a 25-week school implementation reporting improved interoceptive awareness — both without randomisation or adequate power. And it got weaker since: an 8-week online adaptation of that same programme in adults (n = 10) reported “non-significant improvements across all areas”, with the programme’s author as a co-author of the study. The rationale that granularity depends on a person’s OWN concepts, so we should surface theirs rather than impose generic ones, is inference, not a measured effect.
Sources
🎯 Pick for me & Just start
Strength practice consensus
What the feature does
“Pick for me” surfaces one backlog task chosen by a fixed scoring rule — never chance, never AI — to end scrolling a list deciding what to do. “Just start” asks the assistant for the smallest possible first step (2–5 minutes) and opens a calm 2-minute countdown to begin it.
What it rests on
Both halves target a real, well-documented deficit rather than each other. Decision paralysis over an undifferentiated list is choice overload: a meta-analysis of 99 observations (N = 7,202) finds it is real and significant once choice-set complexity, decision difficulty or preference uncertainty are accounted for — and choice deferral (simply not picking) is one of its four interchangeable measures, exactly what “Pick for me” short-circuits by removing the choice. Starting itself is where resistance concentrates: a meta-analysis of 691 correlations identifies task aversiveness and low self-efficacy as the strongest, most consistent predictors of procrastination. Directly on the “smallest step” mechanic, a randomized experiment (N = 1,035) had one group name an easy subtask and a small reward before starting; against neutral-question controls, self-reported completion likelihood rose modestly (d = 0.2, p = .002).
Where the evidence stops
The exact combination we ship — a deterministic single pick, plus an AI-generated 2-minute step with a timer — has not itself been tested by anyone. The closest trial bundles subtask-naming with affect labeling and a chosen reward, measures a single self-reported likelihood item rather than actual completion, and finds a small effect. The famous “2-minute rule” is popular productivity advice (Getting Things Done, Atomic Habits), not a separately trial-tested technique.
Sources
These texts are written in Ukrainian and English. There is no Spanish or French edition: a translated claim about research is a different claim, so we do not make one.