# AWP Planning UX Research Audit

**Date:** 2026-08-19  
**Status:** Research review complete. UXR-01..UXR-07 design direction has been accepted/folded into the Planning baseline; real-user validation remains required before high-fidelity Planning UX is frozen.  
**Scope:** Planning Module product model, Simple/Expert participation, planner-led conversation, Plan Index, Context rail, Quality/CI, delivery/deployment, Connections, readiness, communication behavior, and plan-to-launch handoff.  
**Method:** Heuristic UX-research review using the supplied `ux-researcher` framework: user journeys, behavioral mental models, lean mixed-method validation, usability metrics, drop-off analysis, actionable synthesis, and research ethics.

## 1. Executive finding

The Planning mechanism is structurally strong, but it remains an **expert-designed hypothesis** until representative users successfully use it.

The highest-risk failure mode is:

```text
AWP becomes very good at producing internally coherent plans
without proving that users understand, trust, correct, and successfully act on them.
```

The design now includes a conditional Research & Validation capability and stronger recommendation/progressive-disclosure behavior. The next UX step is a low-fidelity usability study, not another broad layout debate.

## 2. Strong design hypotheses to keep unless evidence disproves them

### F1 — Clear persistent mental model

```text
What did I ask for?     -> Plan Index
What are you doing now? -> Current workspace
What is left / next?    -> planner-led next step
```

This directly reduces memory burden in long AI planning sessions.

### F2 — Planner-led behavior

The planner owns the agenda while users can interrupt, refute, jump, defer and change direction. This is preferable to both a passive chatbot and a rigid wizard.

### F3 — Recommendation-first progressive disclosure

The planner recommends rather than asking users to construct expert answers from scratch. Deeper evidence/options appear when useful.

### F4 — Gate-based readiness

`Ready`, `Ready with accepted/deferred gaps`, and `Blocked` are stronger than fake percentages or exhausting every possible question.

### F5 — Connection interruption recovery

Provider setup preserves/returns to the exact PlanningSession context instead of becoming a separate dead-end workflow.

### F6 — Explicit launch disposition

`Start now / Schedule / Park until later` handles the real case where planning is complete but execution should not begin immediately.

### F7 — Structured state outside chat

Decisions, assumptions, risks, research, evidence, artifacts, connection requirements, technical defaults, execution and delivery policy must remain inspectable without reconstructing a transcript.

## 3. Accepted UX-research amendments

### UXR-01 — Recommendation confidence and basis

Every material recommendation carries structured:

```text
recommendation
confidence: high | medium | low
basis: repository evidence | Project default | accepted decision | industry practice | research | inference
material consequences
what would make AWP recommend differently
```

Progressive disclosure controls density. Low confidence is visible and should not masquerade as certainty.

### UXR-02 — Legible agenda without surrendering planner leadership

Guided planning remains default, but users can see:

```text
Now
Next
Later
```

and may reorder/park a non-blocking item without switching to a fully manual process.

### UXR-03 — Progressive Plan Index

```text
current group expanded
next unresolved group partially visible
completed groups collapsed to summary
blocked/deferred items remain surfaced
search/filter available for large plans
```

A blocker must not disappear because its parent is collapsed.

### UXR-04 — Consequence-based readiness wording

Readiness answers both:

```text
Can I continue now?
What will still block me later?
```

Example:

```text
Ready for specification
You can continue now.
Before production you still need:
  Cloudflare connection
  production approval policy
```

### UXR-05 — Compact integrated Delivery recommendation

Quality + CI + CD + deployment + execution + connections use compact summary rows, with only the active/reviewed dimension expanded.

### UXR-06 — Separate current, intended, desired and recommended state

When they conflict, distinguish:

```text
Observed current state
Documented intended state
User-stated desired state
AWP recommendation
```

Do not collapse these into one inferred truth.

### UXR-07 — Conditional Research & Validation capability

Research is conditional on uncertainty and consequence, not a mandatory stage for every project.

Examples:

```text
new user-facing product + unvalidated assumptions -> user/problem validation
new/high-risk interaction model                  -> low-fi usability validation
existing product with analytics/research         -> inspect evidence first
internal tool for one known team                 -> lightweight/skip may be enough
critical accessibility workflow                  -> representative accessibility validation
```

Research sits conditionally under DEFINE/DESIGN, not as a new universal top-level landmark.

## 4. New participation-mode hypothesis: Simple vs Expert

The Planning UX now separates **planning completeness** from **user involvement depth**.

```text
Simple
  user is asked primarily for OwnerRequired decisions
  AWP resolves DelegableExpert decisions
  technical decisions build visibly and remain editable

Expert
  technical decisions can become active interview items
  user can deeply review/refute recommendations
```

This is a strong response to two different user behaviors:

```text
"I want the product/business outcome; handle the engineering correctly."
"I want to participate deeply in architecture/CI/CD/execution choices."
```

### Research risk R-MODE-1 — Simple could feel opaque

If AWP stops asking technical questions but users cannot easily see what was decided, Simple mode becomes hidden automation.

Required behavior to validate:

```text
technical decisions visibly accumulate
source/default/confidence is readable
changed/low-confidence decisions are surfaced
user can inspect/change them without switching mode
```

### Research risk R-MODE-2 — Simple could still feel technical

If every delegated decision emits a large chat card, Simple mode is not actually simpler.

Use compact `Technical plan updated` feedback and let the Plan Index/technical summary carry persistent detail.

### Research risk R-MODE-3 — Expert may feel like homework

Expert mode must remain recommendation-first. It should not become a long blank questionnaire or assume the user already knows every engineering method.

### Research risk R-MODE-4 — Wrong decision classification

The system must reliably distinguish:

```text
OwnerRequired
DelegableExpert
PolicyRequired
```

The key UX rule is to escalate the **user consequence**, not technical jargon.

### Research risk R-MODE-5 — Defaults can create invisible lock-in

The first Expert Plan may establish Project defaults, but users need to understand what is reusable versus plan-specific.

Project defaults must be explicit, inspectable and versioned; one Plan exception must not silently redefine future Plans.

## 5. Behavioral recruitment segments

Do not invent fictional personas and treat them as evidence. Recruit by actual behavior/responsibility.

Initial segments:

```text
S1  Solo builder / technical founder
S2  Senior/staff engineer or technical lead
S3  Engineering manager / platform / delivery owner
S4  Existing-project maintainer
S5  Less-expert product owner working with coding agents
```

Mode research should deliberately include both users who prefer delegation and users who prefer deep technical control.

## 6. End-to-end Planning journey to validate

```text
J1  Enter/create project
    goal: start useful planning quickly

J2  Choose/understand participation mode
    goal: understand Simple vs Expert without confusing it with project type/profile

J3  Establish intent/scope/users/business decisions
    goal: feel understood without a giant questionnaire

J4  Watch structured/technical plan build
    goal: know what AWP decided without being forced to review everything

J5  Review/correct recommendations
    goal: challenge a recommendation when desired

J6  Resolve/defer provider setup
    goal: connect what is needed without losing planning flow

J7  Resume after interruption
    goal: recover exact context and next step

J8  Review readiness and launch disposition
    goal: understand consequences and choose Start/Schedule/Park confidently

J9  Return to parked/scheduled/active Plan
    goal: understand what changed and what needs attention
```

For existing projects also test conflicts between observed repo state, documentation and desired future state.

## 7. First lean usability study

### Research questions

```text
RQ1  Can users always answer: what did I ask for, what is happening now, what is next?
RQ2  Can users explain Simple vs Expert in their own words?
RQ3  Do Simple users notice technical decisions without feeling forced to review them?
RQ4  Can Simple users find/change one technical decision when they choose?
RQ5  Do Expert users understand recommendation confidence/basis and challenge it appropriately?
RQ6  Can users distinguish inherited/default/inferred/changed technical decisions?
RQ7  Can users distinguish blocking, deferred, accepted-risk and completed items?
RQ8  Does planner-led behavior feel helpful rather than controlling?
RQ9  Can users complete a Connection flow and resume the interrupted decision?
RQ10 Can users resume after interruption without asking what happened?
RQ11 Can users explain what Start, Schedule and Park will do before confirming?
RQ12 Does the three-area workspace reduce memory burden or merely move complexity on screen?
RQ13 Do users understand which Expert decisions are being proposed as future Project defaults?
```

### Core tasks

```text
T1  Start a new software product in Simple mode.
T2  Identify what AWP decided technically without being asked.
T3  Open/change one technical decision.
T4  Correct a wrongly inferred project/profile fact.
T5  Switch to Expert and review one architecture/CI/CD recommendation.
T6  Reject a deployment recommendation and compare alternatives.
T7  Save appropriate Expert decisions as Project defaults; reject a plan-specific one.
T8  Defer a provider connection and identify when it becomes blocking.
T9  Connect a provider and return to the interrupted decision.
T10 Leave/resume the session and identify what was decided and what is next.
T11 Review readiness and choose Schedule or Park.
```

Existing-project variant:

```text
T12 Resolve conflict between repository evidence, documented intent and desired future state.
```

### Metrics

Standard usability measures:

```text
Task success rate
Time on task
Error/recovery rate
Learnability
User confidence/satisfaction
```

Planning-specific:

```text
count of "what next?" / "did we decide this?" recovery questions
Simple -> Expert switch rate and reason
Expert -> Simple switch rate and reason
technical-decision inspection/change success
Project-default comprehension/correction
profile correction rate/time
recommendation override rate + reason
deferred-item resurfacing success
connection return-to-context completion
resume success after >1h / >1d
launch-choice comprehension
```

High conversation length is not success. A shorter session that produces a correct, understood Plan may be better.

### Method

```text
5 representative participants for first round
moderated remote sessions
interactive low-fi prototype or sufficiently functional prototype
30-45 minutes
realistic project data
recording/notes with consent
immediate Finding -> Evidence -> Impact -> Recommendation synthesis
```

Run another small round after material changes rather than one large perfect study.

## 8. Analytics after implementation

Track outcomes, not vanity engagement:

```text
time to first useful structured Plan
planning-stage abandonment
decisions/questions per resolved OwnerRequired item
technical decisions auto-resolved per Plan
technical-decision review/override rate
mode selection/switching
Project-default adoption/change rate
recommendation override rate + reason
reopened Decision rate
deferred item overdue rate
connection setup abandonment + return-to-context
resume success
"what next?" recovery events
launch distribution: start / schedule / park
blocked launch reasons
```

## 9. Research evidence and ethics

If AWP stores interviews, recordings, research participants, surveys or customer research as Planning evidence, treat them as sensitive research data.

Required considerations:

```text
consent/source provenance
participant withdrawal/deletion
PII classification/redaction
retention policy
access control
secure storage
clear separation of research evidence from model-generated assumptions
```

Synthetic personas/model assumptions are never presented as actual research evidence.

## 10. Current disposition

Accepted design direction:

```text
KEEP  planner-led Guided default
KEEP  DEFINE / DESIGN / SPECIFY / DELIVER / LAUNCH landmarks
KEEP  context-following rail
KEEP  recommendation-first alternatives
KEEP  gate-based readiness/deferral
KEEP  Start / Schedule / Park
KEEP  user-global Connections + project-local bindings
KEEP  recommendation-first Quality/CI/CD/execution
ADOPT UXR-01..UXR-07
ADOPT Simple/Expert participation modes
ADOPT visible technical-plan summaries in Simple mode
ADOPT explicit reusable Project-default proposal after Expert planning
```

Still required before high-fidelity Planning mockup freeze:

```text
1. implementation-grade Planning/domain/workflow specs
2. interactive low-fi Planning prototype covering Simple + Expert
3. first lean usability round
4. evidence-backed corrections from that round
```
