Skip to content

Major A* Performance Optimizations (Unified NodeMap, Deferred Allocations, Lean FSA) - #43

Merged
justinhj merged 4 commits into
masterfrom
optimizations-2
Sep 29, 2026
Merged

justinhj merged 4 commits into
masterfrom
optimizations-2

Conversation

@justinhj

@justinhj justinhj commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Summary of Changes

This PR delivers major performance optimizations, container consolidations, and memory allocation reductions for the A* search implementation in stlastar.h and fsa.h.

1. Merged Open & Closed Sets (m_NodeMap)

  • Replaced separate m_OpenSet and m_ClosedList hash containers with a single unified std::unordered_set<Node*, NodeHash, NodeEqual> m_NodeMap.
  • Repurposed the intrusive heap_index on Node as the Open vs. Closed state discriminator:
    • heap_index != SIZE_MAX: Node is currently on the Open min-heap.
    • heap_index == SIZE_MAX: Node is Closed (expanded).
  • Transitioning a node from Open to Closed is now an $O(1)$ assignment (n->heap_index = SIZE_MAX) rather than an erase-from-open plus insert-into-closed.
  • Checking a successor state requires only a single hash lookup instead of two separate queries.

2. Deferred Node Allocation & Inline Streaming Successors

  • Removed the private m_Successors vector entirely.
  • SearchStep() exposes m_CurrentExpandingNode, allowing AddSuccessor() to evaluate neighbors inline as they are generated.
  • A stack-allocated probe (Node dummy; dummy.m_UserState = State;) checks m_NodeMap before allocating memory.
  • AllocateNode() is only invoked when a node is genuinely new and unvisited, completely eliminating allocation churn for the ~75% of neighbors on grid maps that are rejected duplicates.

3. Consistent Heuristic Optimization

  • Added a ConsistentHeuristic = true template parameter (template <class UserState, bool ConsistentHeuristic = true>).
  • For monotonic/consistent heuristics (such as Manhattan or Euclidean distance on grids), closed nodes are mathematically guaranteed to have optimal costs and can never be improved. When ConsistentHeuristic is true, any probe matching a closed node (heap_index == SIZE_MAX) is immediately skipped with zero cost comparisons or re-open overhead.

4. Lean Fixed-Size Block Allocator (fsa.h)

  • Overhauled FixedSizeAllocator to use a lean, singly-linked free list.
  • Removed the unused doubly-linked "used list" tracking (pPrev, pNext, bAllocated), improving cache locality and reducing pointer swaps during alloc() and free().

5. Benchmark & Hash Modernization

  • Updated coordinate hashing in bench.cpp and findpath.cpp to use std::hash<int> instead of std::hash<float>.
  • Benchmark runner scripts and reproducible measurements added to history/ and bench/.

Benchmark Results

1. Unbounded Map-Wide Searches (1,000 x 1,000 Grid, 1,000 searches x 5 runs)

Stage Commit Mean Total Time Std Dev Speedup vs Baseline
Baseline (Pre-optimization) 64bdb20 302.4745 s $\pm 3.7687$ s 1.00x
Step 1: Open set lookup $O(1)$ 7424a4c 86.6529 s $\pm 1.0536$ s 3.49x
Step 2: Intrusive min-heap 3ee685b 20.5033 s $\pm 0.2872$ s 14.75x
Step 3: Unified map & deferred alloc a37013f 13.4744 s $\pm 0.4749$ s 22.44x

Overall: Unbounded full-map search throughput increased by +2144.8% (22.44x faster), cutting runtime by 95.5%.

2. Bounded Tactical Searches (MAX_SEARCH_DISTANCE = 64, 1,000 searches x 5 runs)

Stage Commit Mean Total Time Std Dev Speedup vs Baseline
Baseline (Pre-optimization) 64bdb20 45.3353 s $\pm 1.6622$ s 1.00x
Step 1: Open set lookup $O(1)$ 7424a4c 48.3634 s $\pm 1.7018$ s 0.94x (cache locality wins on tiny frontiers)
Step 2: Intrusive min-heap 3ee685b 44.4163 s $\pm 1.9320$ s 1.02x
Step 3: Unified map & deferred alloc a37013f 23.7850 s $\pm 0.0691$ s 1.91x

Takeaway: While asymptotic gains dominate on large frontiers, eliminating node allocation churn in Step 3 provides a universal ~2x speedup even across small, local searches.


Verification

  • All automated unit tests in tests.cpp pass (ctest --test-dir build --output-on-failure).
  • Example executables (8puzzle, findpath, minpathbucharest) verified functional.

@justinhj justinhj changed the title More optimizations Major A* Performance Optimizations (Unified NodeMap, Deferred Allocations, Lean FSA) Sep 29, 2026
@justinhj
justinhj marked this pull request as ready for review September 29, 2026 04:16
@justinhj
justinhj merged commit 881f636 into master Sep 29, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant