Repository navigation
Support data source sampling with TABLESAMPLE #13563
Description
Activity
I looked around in sqlparser-rs briefly and it seems this syntax is not yet supported
(though the keyword is)
I suggest adding support for this clause in sqlparser first and then we can add a rewrite pass that converts the
TABLESAMPLEclause into a where clause, similar to what we do forSHOW TABLES:datafusion/datafusion/sql/src/statement.rs
Lines 1068 to 1089 in 000288c
fn show_tables_to_plan( &self, extended: bool, full: bool, db_name: Option<Ident>, filter: Option<ShowStatementFilter>, ) -> Result<LogicalPlan> { if self.has_table("information_schema", "tables") { // We only support the basic "SHOW TABLES" // https://github.com/apache/datafusion/issues/3188 if db_name.is_some() || filter.is_some() || full || extended { plan_err!("Unsupported parameters to SHOW TABLES") } else { let query = "SELECT * FROM information_schema.tables;"; let mut rewrite = DFParser::parse_sql(query)?; assert_eq!(rewrite.len(), 1); self.statement_to_plan(rewrite.pop_front().unwrap()) // length of rewrite is 1 } } else { plan_err!("SHOW TABLES is not supported unless information_schema is enabled") } } Thank you for the initial analysis. I will first submit a PR to datafusion-sqlparser-rs with extended grammar.
Reacted by Andrew Lambtake
Sorry for the long absence. After datafusion-sqlparser-rs gained SQL support for tablesamples in 0.54, I am going to introduce a rewrite phase into the datafusion crate.
Reacted by Andrew LambThere are numerous advanced use cases and possible data source-level optimisations for table sampling.
Existing query engines and databases already implement sampling, but it is not in ANSI standard. There are different flavours, but essentially, they allow for specific sampling methods and percentages (or sometimes a number of rows) TABLESAMPLE [SYSTEM | BERNOULLI] (PERCENTAGE | ROWS)
This is my core concern with adding any sort of sampling directly to DataFusion -- I think the usecases will vary widely across systems, and thus I worry that anything we build into DataFusion will likely be fairly complicated as well as not what other systems may want
I think it is actually possible to implement table sampling with the existing APIS through a combination of
- sql planner extension https://github.com/apache/datafusion/blob/main/datafusion-examples/examples/sql_dialect.rs
- User defined extension nodes (aka add extension logical planning nodes)
I would be willing to help make an example for this usecase, to show it is possible. I think it would be a nice showcase for how to extend systems using DataFusion without having to change the ecode.
If there is broad support for adding table sampling to DataFusion, I think we should try and make it conform to the design goals of DataFusion:
- Work “out of the box”: Provide a very fast, world class query engine with minimal setup or required configuration.
- Customizable everything: All behavior should be customizable by implementing traits.
- Architecturally boring 🥱: Follow industrial best practice rather than trying cutting edge, but unproven, techniques.
So maybe in this case we can begin with figuring out "how would we add table sampling"
https://datafusion.apache.org/contributor-guide/architecture.html#creating-new-extension-apisReacted by theirix and Vegard StikbakkeThank you for the detailed explanation, @alamb ! I remember these design goals from community talks. I agree that having table sampling as an extension to DataFusion could be a great way to implement this feature, and at the same time, can showcase the extensibility of the core and identify any gaps with APIs. If this works out, we could avoid the need for that PR to
datafusion-sql.Since the SQL support is already in datafusion-sqlparser-rs, it looks like the main work is on the second part. Let me start exploring what is required to implement this as an extension and get back with any questions.
Reacted by Andrew LambSince the SQL support is already in datafusion-sqlparser-rs, it looks like the main work is on the second part. Let me start exploring what is required to implement this as an extension and get back with any questions.
Thank you @theirix
What I would recommend we do is start working in an example in https://github.com/apache/datafusion/blob/main/datafusion-examples/examples as a way to explore what is possible today. I am happy to take an early look / provide some feedback/ help along the example
I suspect we can do almost everything needed with existing APIs
The high level steps might be:
- Parse the relevant SQL into some sort of
Statement::Sample - Use the LogicalPlanBuilder to create a plan for that new statement
- Implement a custom extension node, if necessary, to implement the sampling
Does that make sense?
- Parse the relevant SQL into some sort of
Since the SQL support is already in datafusion-sqlparser-rs, it looks like the main work is on the second part. Let me start exploring what is required to implement this as an extension and get back with any questions.
I suspect we can do almost everything needed with existing APIs
The high level steps might be:
1. Parse the relevant SQL into some sort of `Statement::Sample` 2. Use the LogicalPlanBuilder to create a plan for that new statement 3. Implement a custom extension node, if necessary, to implement the samplingDoes that make sense?
Yes, absolutely, I am on it. If something is missing from the API, it will be apparent soon.
Reacted by Andrew Lamb- added a commit that references this issue
on Dec 9, 2025
Is your feature request related to a problem or challenge?
It is helpful to have sampling support for queries to ease the exploration of data.
Describe the solution you'd like
It should be supported on the SQL level (
SAMPLEorTABLESAMPLEsyntax). The sampling construct should be passed to the table source so the sampling is performed at the scan plan (e.g. in an optimised parquet reader).This feature could be implemented in three sequential stages:
WHERE RANDOM() < PfilterDescribe alternatives you've considered
It is possible to use
WHERE RANDOM() < 0.1selection (see discussion #13268 ), but the support in SQL is clearer.Existing query engines and databases already implement sampling, but it is not in ANSI standard. There are different flavours, but essentially, they allow for specific sampling methods and percentages (or sometimes a number of rows)
TABLESAMPLE [SYSTEM | BERNOULLI] (PERCENTAGE | ROWS)DuckDB:
PostgreSQL and Trino:
Spark
Clickhouse is different:
Additional context
Also requested in #11554. The filter for sampling was refined in #13268.