Skip to content

[Bug]: SIGSEGV under concurrent fetch + insert when the writer crosses a segment switch #714

Description

@YongqiYin

Description

One writer thread running insert() concurrently with reader threads running fetch() / query() / group_by_query() crashes the process with SIGSEGV on Linux when the writer crosses a segment-switch boundary. Reproduced **3/3 on main .

Root cause. Readers snapshot the segment list via get_all_segments() with no lock, while writers mutate that state under exclusive write_mtx_ ,doc_ids_ grows per insert , and writing_segment_ is reassigned per switch. Worse, dump()/flush() tear down the old segment's in-memory store (memory_store_.reset()) without taking seg_mtx_ at all , so a reader mid-Fetch — even holding seg_mtx_ — reads a store that is being destroyed.

gdb captures both threads in the same instant (full stacks below):

  • Reader: fetch → SegmentImpl::Fetch → MemForwardStore::convertToTable → arrow::Table::FromRecordBatches → SIGSEGV
  • Writer: write_impl → switch_to_new_segment_for_writing → dump() → flush() → LocalWalFile::remove()

On macOS the same race does not crash; a probe (150µs widened reader window) measured 99.7% of doc_ids_ push_backs and 99.4% of segment-switch reassignments executing while a reader was inside get_all_segments(). The unsynchronized shape dates to the initial commit.

Why rarely hit: only the doc-count-threshold switch is exposed — the other seven switch call sites (optimize, DDL, iterator creation) hold the schema lock exclusively and drain in-flight readers first (verified: 285 optimize-triggered switches, zero crashes). But the threshold counts cumulative inserts, not collection size (doc_ids_ is append-only), so any long-running writer eventually crosses it.

Steps to Reproduce

1. Build zvec on Linux x86_64 (glibc), `CMAKE_BUILD_TYPE=Release`.
2. Collection with a vector field; `max_doc_count_per_segment` = 4,000 (schema-validated: `validate()` accepts ≥ 1,000) and 8 MB forward buffer — this compresses the switch interval from ~10M inserts to seconds without changing the code path.
3. Preload ~3,000 docs.
4. One writer thread: `insert()` batches of 100 in a tight loop.
5. Four reader threads: `fetch()` preloaded docs by pk in a tight loop.
6. SIGSEGV within seconds.

Logs / Stack Trace

Thread (reader) — SIGSEGV:
#0  arrow::Field::ToString() const
#1  arrow::Schema::ToString() const
#2  arrow::Table::FromRecordBatches(...)
#3  zvec::MemForwardStore::convertToTable(...)
#4  zvec::MemForwardStore::fetch(...)
#5  zvec::SegmentImpl::fetch_normal(...)
#6  zvec::SegmentImpl::Fetch(unsigned long, ...)
#7  zvec::CollectionImpl::fetch(...)

Concurrent writer thread (same instant):
#2 std::filesystem::remove(...)
#3 zvec::ailego::FileHelper::DeleteFile(char const*)
#4 zvec::LocalWalFile::remove()
#5 zvec::SegmentImpl::flush()
#6 zvec::SegmentImpl::dump()
#7 zvec::CollectionImpl::switch_to_new_segment_for_writing(...)

Operating System

Linux x86_64 (Alibaba Cloud ECS, 64-core Xeon Platinum 8369B, glibc 2.32, GCC 10.2.1). Not observed to crash on macOS arm64 under the same load (race still active there; see Description).

Build & Runtime Environment

zvec main @ 3d479a1 (also reproduced at 733903c). CMake + Ninja, Release, C++17.

Additional Context

  • I've checked git status — no uncommitted submodule changes
  • I built with CMAKE_BUILD_TYPE=Debug
  • This occurs with or without COVERAGE=ON
  • The issue involves Python ↔ C++ integration (pybind11)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

bugSomething isn't working

Type

No type

Projects

  • Status
    Done

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions