With the spooling protocol, the coordinator writes result segments to object storage and deletes each segment when the client acknowledges it (fs.segment.explicit-ack=true, the default).
SegmentIterator acknowledges a spooled segment only when the caller asks for the row after that segment's last row. A read that stops at the last row never acknowledges the final segment:
cursor.fetchone() on a single-row result
- SQLAlchemy
.scalar(), .first(), .one()
- the SQLAlchemy dialect's own connect-time queries (with
protocol.spooling.inlining.enabled=false)
That segment stays in the spooling location until the server's expired-segment pruning deletes it (fs.segment.ttl, 12 hours by default). So small query results stay in the bucket long after the client has read them.
Reproduced on 0.340.0 and master 74c6156 against Trino 483, with a filesystem spooling manager on AWS S3, retrieval modes coordinator_proxy and storage, and inlining disabled:
import sqlalchemy as sa
engine = sa.create_engine("trino://user@localhost:8080/tpch")
with engine.connect() as conn:
conn.execute(sa.text("SELECT count(*) FROM tpch.tiny.nation")).scalar()
# listing the spooling prefix afterwards: 5 segment objects remain
# (the connect-time queries plus the scalar() query)
with engine.connect() as conn:
conn.execute(sa.text("SELECT name FROM tpch.tiny.nation")).all()
# .all() reads past the last row, so its segment is acknowledged and deleted
A decoded segment's rows are already in memory, so it can be acknowledged as soon as decoding succeeds. A segment whose download or decode fails would still not be acknowledged, and is retried as it is today.
With the spooling protocol, the coordinator writes result segments to object storage and deletes each segment when the client acknowledges it (
fs.segment.explicit-ack=true, the default).SegmentIteratoracknowledges a spooled segment only when the caller asks for the row after that segment's last row. A read that stops at the last row never acknowledges the final segment:cursor.fetchone()on a single-row result.scalar(),.first(),.one()protocol.spooling.inlining.enabled=false)That segment stays in the spooling location until the server's expired-segment pruning deletes it (
fs.segment.ttl, 12 hours by default). So small query results stay in the bucket long after the client has read them.Reproduced on 0.340.0 and master 74c6156 against Trino 483, with a filesystem spooling manager on AWS S3, retrieval modes
coordinator_proxyandstorage, and inlining disabled:A decoded segment's rows are already in memory, so it can be acknowledged as soon as decoding succeeds. A segment whose download or decode fails would still not be acknowledged, and is retried as it is today.