Major A* Performance Optimizations (Unified NodeMap, Deferred Allocations, Lean FSA) - #43
Merged
Merged
Conversation
justinhj
marked this pull request as ready for review
September 29, 2026 04:16
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary of Changes
This PR delivers major performance optimizations, container consolidations, and memory allocation reductions for the A* search implementation in
stlastar.handfsa.h.1. Merged Open & Closed Sets (
m_NodeMap)m_OpenSetandm_ClosedListhash containers with a single unifiedstd::unordered_set<Node*, NodeHash, NodeEqual> m_NodeMap.heap_indexonNodeas the Open vs. Closed state discriminator:heap_index != SIZE_MAX: Node is currently on the Open min-heap.heap_index == SIZE_MAX: Node is Closed (expanded).n->heap_index = SIZE_MAX) rather than an erase-from-open plus insert-into-closed.2. Deferred Node Allocation & Inline Streaming Successors
m_Successorsvector entirely.SearchStep()exposesm_CurrentExpandingNode, allowingAddSuccessor()to evaluate neighbors inline as they are generated.Node dummy; dummy.m_UserState = State;) checksm_NodeMapbefore allocating memory.AllocateNode()is only invoked when a node is genuinely new and unvisited, completely eliminating allocation churn for the ~75% of neighbors on grid maps that are rejected duplicates.3. Consistent Heuristic Optimization
ConsistentHeuristic = truetemplate parameter (template <class UserState, bool ConsistentHeuristic = true>).ConsistentHeuristicistrue, any probe matching a closed node (heap_index == SIZE_MAX) is immediately skipped with zero cost comparisons or re-open overhead.4. Lean Fixed-Size Block Allocator (
fsa.h)FixedSizeAllocatorto use a lean, singly-linked free list.pPrev,pNext,bAllocated), improving cache locality and reducing pointer swaps duringalloc()andfree().5. Benchmark & Hash Modernization
bench.cppandfindpath.cppto usestd::hash<int>instead ofstd::hash<float>.history/andbench/.Benchmark Results
1. Unbounded Map-Wide Searches (1,000 x 1,000 Grid, 1,000 searches x 5 runs)
64bdb207424a4c3ee685ba37013f2. Bounded Tactical Searches (
MAX_SEARCH_DISTANCE = 64, 1,000 searches x 5 runs)64bdb207424a4c3ee685ba37013fVerification
tests.cpppass (ctest --test-dir build --output-on-failure).8puzzle,findpath,minpathbucharest) verified functional.