Skip to content

Support data source sampling with TABLESAMPLE #13563

Description

@theirix

Is your feature request related to a problem or challenge?

It is helpful to have sampling support for queries to ease the exploration of data.

Describe the solution you'd like

It should be supported on the SQL level (SAMPLE or TABLESAMPLE syntax). The sampling construct should be passed to the table source so the sampling is performed at the scan plan (e.g. in an optimised parquet reader).

This feature could be implemented in three sequential stages:

  1. Support additional SQL syntax but fail in the physical plan builder
  2. Transparently convert to WHERE RANDOM() < P filter
  3. For eligible data sources push the sampling to the table source

Describe alternatives you've considered

It is possible to use WHERE RANDOM() < 0.1 selection (see discussion #13268 ), but the support in SQL is clearer.

Existing query engines and databases already implement sampling, but it is not in ANSI standard. There are different flavours, but essentially, they allow for specific sampling methods and percentages (or sometimes a number of rows) TABLESAMPLE [SYSTEM | BERNOULLI] (PERCENTAGE | ROWS)

DuckDB:

SELECT * FROM tbl TABLESAMPLE SYSTEM (10%),

PostgreSQL and Trino:

SELECT * FROM tbl TABLESAMPLE SYSTEM (10),

Spark

SELECT * FROM tbl TABLESAMPLE SYSTEM (10 PERCENT)

Clickhouse is different:

SELECT * FROM tbl SAMPLE 0.1

Additional context

Also requested in #11554. The filter for sampling was refined in #13268.

Activity

  1. alamb commented on Nov 25, 2024

    @alamb
    Contributor

    I looked around in sqlparser-rs briefly and it seems this syntax is not yet supported

    https://github.com/search?q=repo%3Aapache%2Fdatafusion-sqlparser-rs%20tablesample&type=code

    (though the keyword is)

    I suggest adding support for this clause in sqlparser first and then we can add a rewrite pass that converts the TABLESAMPLE clause into a where clause, similar to what we do for SHOW TABLES:

    fn show_tables_to_plan(
    &self,
    extended: bool,
    full: bool,
    db_name: Option<Ident>,
    filter: Option<ShowStatementFilter>,
    ) -> Result<LogicalPlan> {
    if self.has_table("information_schema", "tables") {
    // We only support the basic "SHOW TABLES"
    // https://github.com/apache/datafusion/issues/3188
    if db_name.is_some() || filter.is_some() || full || extended {
    plan_err!("Unsupported parameters to SHOW TABLES")
    } else {
    let query = "SELECT * FROM information_schema.tables;";
    let mut rewrite = DFParser::parse_sql(query)?;
    assert_eq!(rewrite.len(), 1);
    self.statement_to_plan(rewrite.pop_front().unwrap()) // length of rewrite is 1
    }
    } else {
    plan_err!("SHOW TABLES is not supported unless information_schema is enabled")
    }
    }

  2. theirix commented on Nov 25, 2024

    @theirix
    ContributorAuthor

    Thank you for the initial analysis. I will first submit a PR to datafusion-sqlparser-rs with extended grammar.

  3. theirix commented on Nov 25, 2024

    @theirix
    ContributorAuthor

    take

  4. theirix commented on May 19, 2025

    @theirix
    ContributorAuthor

    Sorry for the long absence. After datafusion-sqlparser-rs gained SQL support for tablesamples in 0.54, I am going to introduce a rewrite phase into the datafusion crate.

  5. alamb commented on Aug 18, 2025

    @alamb
    Contributor

    There are numerous advanced use cases and possible data source-level optimisations for table sampling.

    Existing query engines and databases already implement sampling, but it is not in ANSI standard. There are different flavours, but essentially, they allow for specific sampling methods and percentages (or sometimes a number of rows) TABLESAMPLE [SYSTEM | BERNOULLI] (PERCENTAGE | ROWS)

    This is my core concern with adding any sort of sampling directly to DataFusion -- I think the usecases will vary widely across systems, and thus I worry that anything we build into DataFusion will likely be fairly complicated as well as not what other systems may want

    I think it is actually possible to implement table sampling with the existing APIS through a combination of

    1. sql planner extension https://github.com/apache/datafusion/blob/main/datafusion-examples/examples/sql_dialect.rs
    2. User defined extension nodes (aka add extension logical planning nodes)

    I would be willing to help make an example for this usecase, to show it is possible. I think it would be a nice showcase for how to extend systems using DataFusion without having to change the ecode.

    If there is broad support for adding table sampling to DataFusion, I think we should try and make it conform to the design goals of DataFusion:

    1. Work “out of the box”: Provide a very fast, world class query engine with minimal setup or required configuration.
    2. Customizable everything: All behavior should be customizable by implementing traits.
    3. Architecturally boring 🥱: Follow industrial best practice rather than trying cutting edge, but unproven, techniques.

    So maybe in this case we can begin with figuring out "how would we add table sampling"
    https://datafusion.apache.org/contributor-guide/architecture.html#creating-new-extension-apis

  6. theirix commented on Aug 18, 2025

    @theirix
    ContributorAuthor

    Thank you for the detailed explanation, @alamb ! I remember these design goals from community talks. I agree that having table sampling as an extension to DataFusion could be a great way to implement this feature, and at the same time, can showcase the extensibility of the core and identify any gaps with APIs. If this works out, we could avoid the need for that PR to datafusion-sql.

    Since the SQL support is already in datafusion-sqlparser-rs, it looks like the main work is on the second part. Let me start exploring what is required to implement this as an extension and get back with any questions.

  7. alamb commented on Aug 19, 2025

    @alamb
    Contributor

    Since the SQL support is already in datafusion-sqlparser-rs, it looks like the main work is on the second part. Let me start exploring what is required to implement this as an extension and get back with any questions.

    Thank you @theirix

    What I would recommend we do is start working in an example in https://github.com/apache/datafusion/blob/main/datafusion-examples/examples as a way to explore what is possible today. I am happy to take an early look / provide some feedback/ help along the example

    I suspect we can do almost everything needed with existing APIs

    The high level steps might be:

    1. Parse the relevant SQL into some sort of Statement::Sample
    2. Use the LogicalPlanBuilder to create a plan for that new statement
    3. Implement a custom extension node, if necessary, to implement the sampling

    Does that make sense?

  8. theirix commented on Aug 20, 2025

    @theirix
    ContributorAuthor

    Since the SQL support is already in datafusion-sqlparser-rs, it looks like the main work is on the second part. Let me start exploring what is required to implement this as an extension and get back with any questions.

    I suspect we can do almost everything needed with existing APIs

    The high level steps might be:

    1. Parse the relevant SQL into some sort of `Statement::Sample`
    
    2. Use the LogicalPlanBuilder to create a plan for that new statement
    
    3. Implement a custom extension node, if necessary, to implement the sampling
    

    Does that make sense?

    Yes, absolutely, I am on it. If something is missing from the API, it will be apparent soon.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions