According to Fig. 12(b), the PQ computation time includes top-k token index time, without the communication latency of top-k tokens' kv from host CPU. However, in the pq_search.py file, the pq_start and pq_end functions profile the entire process including the communication latency (cache_managers[self.rank].fetch_and_concat_kv_w_cache), which contradicts what the paper said. Could you please explain on this?
According to Fig. 12(b), the PQ computation time includes top-k token index time, without the communication latency of top-k tokens' kv from host CPU. However, in the pq_search.py file, the pq_start and pq_end functions profile the entire process including the communication latency (cache_managers[self.rank].fetch_and_concat_kv_w_cache), which contradicts what the paper said. Could you please explain on this?