Skip to content

v0.1.0 release hardening plan #1

Description

@zekageri

Goal

Make Phase safe and verifiable for the first public v0.1.0 release.

This plan is based on the release review of main at commit 00b6ac9a6b4660f83c2547795d84e6421ad0414e.

Release blockers

1. Make task teardown ownership race-free

The worker currently clears taskRunning and taskHandle before it has actually finished executing and deleted/suspended itself. Phase::end() may therefore return and allow _impl destruction while the worker still accesses PhaseImpl.

  • Redesign worker termination so end() cannot observe completion before the task has stopped accessing PhaseImpl.
  • Prefer an explicit worker-exit handshake:
    • worker completes shutdown and final diagnostics
    • worker signals exitReady
    • worker suspends or waits without touching PhaseImpl
    • caller-side end() deletes the worker task using the stored handle
    • only then clear the handle and mark the object ended
  • Support both normal FreeRTOS tasks and capability-created PSRAM-stack tasks.
  • Ensure task memory is released exactly once.
  • Ensure semaphore/mutex destruction only happens after worker termination.
  • Add deterministic tests for end() from Idle, Booting, Starting, Ready, Paused, Stopping, Failed, and Stopped.
  • Add a repeated create/init/start/end/destroy stress test.

2. Reject end() from the Phase task

Callbacks execute synchronously on the Phase task. Calling end() from a lifecycle callback, group condition, onChange, onReady, or onFailed can make the task wait for itself.

  • Detect xTaskGetCurrentTaskHandle() == taskHandle in end().
  • Return a clear non-success result without blocking.
  • Document that end() must be called by another task.
  • Define destructor behavior when destruction is attempted from the Phase task.
  • Add tests covering end() from every callback type.
  • Verify that stop(), pause(), and resume() remain safe when called from callbacks.

3. Fix callback string lifetime guarantees

PhaseChange::pauseReason points into mutable internal storage. Another task can call resume() while the callback is running and invalidate the pointer.

  • Snapshot pauseReason into storage whose lifetime covers the complete callback invocation.
  • Review nodeName and message for the same guarantee.
  • Keep callbacks outside the internal mutex.
  • Add a concurrency test where another task changes pause state while onChange is executing.
  • Keep the documented guarantee that pointers are valid for the duration of the callback.

4. Remove lifecycle-time dynamic allocations

Runtime snapshot objects currently copy std::string, dependency vectors, and std::function instances during boot/readiness/shutdown.

  • Make the registered graph immutable after registration closes.
  • Replace allocating runtime snapshots with index-based or fixed lightweight snapshots.
  • Avoid copying dependency vectors during lifecycle evaluation.
  • Avoid copying node names during normal lifecycle execution where a stable immutable reference/index is sufficient.
  • Avoid copying std::function unless the implementation proves the copy is non-allocating and safe.
  • Ensure no lifecycle path can terminate because a snapshot allocation failed.
  • Update documentation to accurately distinguish setup-time allocation from runtime behavior.
  • Add an allocation audit/test around start, group polling, stop, rollback, and end.

5. Add behavioral test coverage

Current CI compiles examples across multiple ESP32 targets but does not validate state-machine behavior.

  • Add deterministic host tests or an ESP32 test harness for:
    • dependency-ordered init
    • dependency-ordered start/readiness
    • reverse stop order
    • reverse deinit order
    • required init failure rollback
    • required start failure rollback
    • optional init failure
    • optional start failure and deinit cleanup
    • optional dependent skipping
    • missing dependency validation
    • circular dependency validation
    • group success
    • group timeout
    • pause/resume during group polling
    • stop immediately after start()
    • stop after the worker consumes the start request
    • repeated stop requests
    • restart after Stopped
    • concurrent diagnostics/state reads
    • callback-triggered stop/pause/resume
    • callback-triggered end() rejection
    • teardown synchronization
  • Make behavioral tests blocking in CI.
  • Keep multi-board example builds as compatibility checks.

6. Align release metadata with the tag

  • Set library.json version to 0.1.0.
  • Set library.properties version to 0.1.0.
  • Update README status from 0.0.1 to 0.1.0.
  • Add a blocking CI check that vX.Y.Z matches both manifests.
  • Run Arduino metadata lint as a blocking release prerequisite, or add an equivalent strict metadata check.

Correctness and observability fixes

7. Correct stack high-water-mark units

ESP-IDF reports uxTaskGetStackHighWaterMark() in bytes. Phase currently multiplies it by sizeof(StackType_t).

  • Return the value without multiplying it.
  • Add a focused test or platform assertion for the diagnostic unit.
  • Confirm documentation says bytes.

8. Define shutdown failure semantics

Stop/deinit callback failures are currently emitted but discarded while shutdown still ends in Stopped.

Choose and document one policy:

  • Best effort: continue shutdown, aggregate failures, expose the final shutdown result and diagnostics.

or

  • Event-only: continue shutdown and explicitly document that stop/deinit failures are observable only through onChange().

Also:

  • Preserve cleanup progress even when one stop/deinit callback fails.
  • Add tests with multiple failing shutdown callbacks.

API and documentation pass

  • Document exact async semantics of start(), stop(), pause(), and resume().
  • Document allowed state transitions and restart behavior after Stopped.
  • Document that callback timeouts are post-return measurements and cannot interrupt stuck callbacks.
  • Document callback execution context and reentrancy rules.
  • Document end() as terminal and external-task-only.
  • Document allocation behavior after runtime-allocation changes.
  • Verify every example checks registration builder results where practical.
  • Add a release changelog summarizing the supported API and known cooperative limitations.

Recommended implementation order

  1. Teardown ownership and worker-exit handshake.
  2. Self-end()/destructor policy.
  3. Callback pointer lifetime fixes.
  4. Runtime snapshot/allocation redesign.
  5. Behavioral test harness and race tests.
  6. Stack diagnostic and shutdown-result fixes.
  7. Metadata, documentation, and release validation.

Release acceptance criteria

v0.1.0 is ready when:

  • No task can access PhaseImpl after end() returns.
  • end() cannot deadlock when invoked from the Phase task.
  • All PhaseChange pointers remain valid throughout callback execution.
  • Lifecycle execution does not depend on unhandled dynamic allocations.
  • Behavioral tests cover success, failure, rollback, pause, cancellation, restart, and teardown races.
  • CI is green for behavioral tests and all supported board builds.
  • Release tag and manifest versions match 0.1.0.
  • Documentation matches the implemented lifecycle and shutdown semantics.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions