Skip to content

[DISCUSS] DataFusion less frequent major / breaking releases (ease using multiple third-party extensions (like delta, or iceberg) ) #16622

Description

@alamb

Is your feature request related to a problem or challenge?

One of the dreams of the composable data ecosystem is to quickly assemble a system from various components (DataFusion, data formats

DataFusion still releases once a month, which allows code to quickly flow but also causes at least 2 challenges:

  1. Takes non trivial work required to upgrade downstream projects, as mentioned in [Discuss] Release cadence / patch releases / Long Term Supported (lts) minor releases #5269
  2. Make upgrading and using downstream third-party extensions hard

Third party extensions like delta-rs and iceberg provide TableProviders for DataFusion, which is really nice. However, to use those packages the versions of DataFusion must match exactly.

This means for an application that relies on multiple downstream packages must wait until ALL of them have upgraded to the new version in order to upgrade DataFusion. If there is any delay in the downstream libraries updating, it delays.

For example, an application that wants to use delta-rs, iceberg, and the table-providers crate, there is a race after each upgrade of DataFusion

Let's take a release timeline for

  1. +0 days: DataFusion version X released
  2. +7 days: New delta-rs releases upgraded to DataFusion X
  3. +11 days: new iceberg crate released upgraded to DataFusion X
  4. +12 days: new table-providers version is released
  5. +13-30 days: End user app can upgrade DataFusion and delta, and icerberg
  6. +31 days: New DataFusion is released again

Describe the solution you'd like

I would like downstream libraries to have more time and schedule flexibility when upgrading DataFusion and other dependent crates, so that it is easier to construct a system from different components

Describe alternatives you've considered

Option 1: Switch to major/minor release cadence

We could follow the model of arrow-rs which does releases monthly, but breaking releases only quarterly. Here is how it works in arrow-rs: https://github.com/apache/arrow-rs?tab=readme-ov-file#release-versioning-and-schedule

This would mean continuing to release every month, but only allowing breaking API changes every 3rd release (or some other cadence)

The major cost here is that maintainers and contributors would have to be diligent about not merging breaking API changes until a major release

This is possible to automate somewhat:

Option 2: LTS and feature branch

-Keep (at least) two branches going: LTS and main, as proposed by @andygrove in #5269

In this model we would likely backport changes to the LTS branch and make releases from there. The downside of this approach is that there is extra work to backport changes to LTS.

Option 3: Split API into Stable and Unstable parts

This is described by @thinkharderdev in

The idea is to extract the API into separate crates (e.g. datafusion-catalog-api) that don't change as frequently and where the changes are more carefully controlled

Additional context

No response

Activity

  1. changed the title [-][DISCUSS] DataFusion minor releases (ease using multiple third-party extensions (like delta, or iceberg) )[/-] [+][DISCUSS] DataFusion less frequent major / breaking releases (ease using multiple third-party extensions (like delta, or iceberg) )[/+] on Jun 30, 2025
  2. findepi commented on Jun 30, 2025

    @findepi
    Member

    I would like downstream libraries to have more time and schedule flexibility when upgrading DataFusion and other dependent crates, so that it is easier to construct a system from different components

    That's a noble goal.

    At the same time we should avoid policies that slow down DataFusion development.

    Perhaps part of a solution could be to invert dependencies between the project. For example, Iceberg-DataFusion and Delta-DataFusion integrations could live in this repo.

  3. alamb commented on Jun 30, 2025

    @alamb
    ContributorAuthor

    Perhaps part of a solution could be to invert dependencies between the project. For example, Iceberg-DataFusion and Delta-DataFusion integrations could live in this repo.

    That would make running tests in the DataFusion repo harder, perhaps (certainly there would be more of them).

    Also, while it could work for icerberg and delta-rs potentially, I am not sure we could plausibly bring all the integrations into the same repo (though I will admit it seems to work for trino 🤔 )

    I am not opposed necessairly, just trying to think of the trade offs

    Another approach maybe could be to break the datafusion repo into separate repos with less stuff that can move slower and keep as much as possible in other repos that can move faster 🤔

  4. milenkovicm commented on Jun 30, 2025

    @milenkovicm
    Contributor

    IMHO, option 1 may be better, it is not as frequent as current release cycle but still frequent enough to take breaking changes in relatively small and digestible chunks.

    Not sure if option 2 would force downstream project to have same branching strategy. Otherwise it would bring "big update" every time LTS was released.

  5. alamb commented on Jun 30, 2025

    @alamb
    ContributorAuthor

    And to be clear, in my mind option 1 still has us doing monthly releases, we just restrict which releases have breaking API changes

  6. timsaucer commented on Jun 30, 2025

    @timsaucer
    Member

    I tend to favor option 1. I am currently in this same position where each version of DF requires updates to Lance and datafusion-python. In the last update, within a week of getting my repo upgraded to DF47, DF48 was released.

    The part that is unclear to me, but maybe arrow-rs has already solved, is how we go about going PR reviews on work that does introduce breaking changes. Consider the timeline where a feature is developed and ready for review two weeks after a major release. That feature contains breaking API changes. Do we then require that feature to sit in a long living PR for another 6-8 weeks until we merge it to main? Those can be extremely difficult to maintain.

  7. findepi commented on Jun 30, 2025

    @findepi
    Member

    If we look at the problem through the lens of project consumers, the expectation to never break anything is natural. If we look at the problem of feature implementers, the exception to be able to introduce changes, even those that are unfortunately breaking, is natural.

    arrow-rs has 2.6x less code and conceptually much simpler API surface. DataFusion is a library for building query engines and data processing tools. It is in active development and is not yet in "mostly complete, mostly maintenance" stage, so internal API changes are often unavoidable. DataFusion being a library, it has enormous API surface, so most of "internal API changes" are actually potentially breaking changes for downstream consumers.

    Being able to do such changes once a quarter (4 times a year), looks too rare to me.

    Also, while it could work for icerberg and delta-rs potentially, I am not sure we could plausibly bring all the integrations into the same repo

    Agreed. At the same time, this issue is not on context of any integrations. It's in a context of real-life problems we (some of us) have today. What are these problems, and what integrations they pertain?

  8. alamb commented on Jun 30, 2025

    @alamb
    ContributorAuthor

    This is a pretty sweet idea from @jonmmease about making upgrades easier (use LLM agents): #13648 (comment)

    (it is somewhat orthogonal to the other goals of this ticket, but I thought it was so good I wanted to share. Maybe like all great ideas it seems obvious in retrospect)

  9. andygrove commented on Jun 30, 2025

    @andygrove
    Member

    I favor option 2 (as already pointed out in the original issue). I don't think that we should artificially slow down development against the main branch.

    We already have a documented process for backporting PRs to release branches (e.g. branch-48) and creating future 48.x.x releases without breaking changes. We always release from the release branches, not the main branch. This is the same model that Apache Spark uses. We just haven't been proactive in backporting PRs to the release branches most of the time.

    The idea of holding off on merging PRs to main sounds like a nightmare to me. In Comet, we often temporarily point to a fixed revision of DataFusion's main branch so that we can gradually upgrade to the next version, testing all recently merged PRs. If breaking change PRs are all sitting in a queue until just before the next major DataFusion release, then I'm not sure how we can gradually upgrade.

  10. andygrove commented on Jun 30, 2025

    @andygrove
    Member

    The responsibility for creating the PRs to backport fixes to the release branch should fall to the downstream users who are waiting on those fixes. There is additional work for DataFusion maintainers in terms of PR reviews though.

  11. comphead commented on Jun 30, 2025

    @comphead
    Contributor

    More inclining to version 2 with slight modifications. Similar to Rustc flow https://web.mit.edu/rust-lang_v1.25/arch/amd64_ubuntu1404/share/doc/rust/html/book/second-edition/ch01-03-how-rust-is-made-and-nightly-rust.html

    so it would be a main branch representing latest update(keep in mind main can be broken by bad commits, performance issue, etc)

    a nightly release branch where DF releases artifacts per night or per week, this branch should be protected by extended tests, performance tests, etc

  12. alamb commented on Jun 30, 2025

    @alamb
    ContributorAuthor

    So for Version 2 (a LTS branch) do we have any proposal for the cadence of releases?

    Like would we do 2 releases each month now?

    1. LTS release
    2. Major release from main

    Or would we perhaps keep one release a month and release on alternate months 🤔

  13. djouallah commented on Jul 1, 2025

    @djouallah

    if I may, duckdb process seems a bit reasonable, release a major version with breaking change every three months (1.1, 1.2 , 1.3 etc), release a bug fix every month, 1.3.1 , 1.3.2 then next version will be 1.4, and they have nightly build for both the main branch and the stable branch

  14. alamb commented on Jul 1, 2025

    @alamb
    ContributorAuthor

    if I may, duckdb process seems a bit reasonable, release a major version with breaking change every three months (1.1, 1.2 , 1.3 etc), release a bug fix every month, 1.3.1 , 1.3.2 then next version will be 1.4, and they have nightly build for both the main branch and the stable branch

    This is basically the same cadence that we use in arrow-rs as well. The only potential issue I see with doing something like that for DataFusion is accumulating 3 months of breaking changes might be a lot 🤔

  15. 14 remaining items

  16. milenkovicm commented on Sep 12, 2026

    @milenkovicm
    Contributor

    Can we consider "breaking change per quarter budget", it would raise a bar for merging braking changes,

    Wouldn't this directly interfere with the comment above:

    I believe major releases every six to eight weeks helped datafusion growth

    Meaning how much of the growth are we leaving on table because we ran out of breaking change budget? Also what would be the criteria for accepting/rejecting a breaking changes and how do we track them.

    We do not know what the percentage of breaking changes were made as there were no other options and what percentage ware done as that was the easy path forward, having some kind of budget could create a pushback on later.

    Criteria for accepting breaking change was deliberately left out for discussion, it could be number of pmc or mainintainer endorsement or similar

  17. andygrove commented on Sep 12, 2026

    @andygrove
    Member
  18. Omega359 commented on Sep 12, 2026

    @Omega359
    Contributor

    In some ways our organization is already doing this - I typically update our internal datafusion version every 3rd or 4th release as I just don't have the bandwidth to test every single release. I'd love a LTS version for that reason if the release cycle for that was every 6-12 months.

  19. kumarUjjawal commented on Sep 12, 2026

    @kumarUjjawal
    Contributor

    I'd love a LTS version for that reason if the release cycle for that was every 6-12 months.

    I feel the same. Even though there would be a lot of work involved for the people involved in the release process. It would be a good step to have TLS.

  20. nuno-faria commented on Sep 12, 2026

    @nuno-faria
    Contributor

    Somewhat related to this, I think it would be nice if a merged PR had a way to quickly tell which version it applies to (or will apply). Sometimes in the past I had to check the changelog to confirm if a PR got in the current release or not. We could use milestones (like Rust or Kubernetes), or regular labels (like Helm):

    Image Image

    This should be easy to do with actions that run on merge.

    If then DataFusion supports an LTS version, we could add a label there as well so it is easy to tell that it was backported.

  21. alamb commented on Sep 13, 2026

    @alamb
    ContributorAuthor

    @alamb High-level, an LTS in some form could definitely make sense. Coming from the more traditional database world, the fact that we have thus far gotten away with shipping API-breaking releases every 6-8 weeks is a little surprising 😅

    Indeed!

    1. Six months is a pretty short support duration for an LTS release. Is that long enough that it will sufficiently ease downstream pain?

    I think it would ease downstream pain compared to 1.5 months we have now 😅 -- but for sure the longer we supported releases the easier it would be for downstream projects to sync.

    1. What should the criteria be for backporting changes to the LTS branch?

    I would initially propose a pretty lenient policy (basically anything that doesn't have API changes and is not deemed "too risky" by maintainers). For example, we probably wouldn't backport refactors or major new features like AS OF Joins. However, if we find we are introducing too many regressions we could tighten up the backport policy.

    But maybe coding agents make it more feasible to do aggressive backporting?

    This is my thesis and I think we could try it for a while

    1. Should we consider shipping DataFusion releases less often?

    This is also a possibility (though I would personally suggest we go to the major/minor release model described in this issue's description as it has worked well for arrow). I think it would work just as well, but would shift the burden as it would restrict development on main

    I typically update our internal datafusion version every 3rd or 4th release as I just don't have the bandwidth to test every single release. I'd love a LTS version for that reason if the release cycle for that was every 6-12 months.

    @Omega359 we have basically done this at InfluxData too (not deliberately, but this is what has sort of happened organically)

  22. alamb commented on Sep 13, 2026

    @alamb
    ContributorAuthor

    I think it would be nice if a merged PR had a way to quickly tell which version it applies to (or will apply). Sometimes in the past I had to check the changelog to confirm if a PR got in the current release or not. We could use milestones (like Rust or Kubernetes), or regular labels (like Helm):

    @nuno-faria I agree this would be super helpful and is not clear from PRs at the moment. Maybe it is worth starting another issue on this topic 🤔

  23. alamb commented on Sep 13, 2026

    @alamb
    ContributorAuthor

    Can we consider "breaking change per quarter budget", it would raise a bar for merging braking changes, yet it would allow controlled number of braking change to get in.

    @milenkovicm this is a very interesting idea and one I think we should pursue in parallel -- maybe we could start with a "really don't break the API unless needed" type wording for PRs (aka really emphasize using deprecation rather than changing signatures for example)

  24. gruuya commented on Sep 17, 2026

    @gruuya
    Contributor

    I noticed the idea of DF LTS versions being discussed in the context of the DF 56.0.0 release. As someone who's gone through many many DF updates (we haven't missed a single one since DF 9.0.0!), I'd like to share my comments.

    This means for an application that relies on multiple downstream packages must wait until ALL of them have upgraded to the new version in order to upgrade DataFusion. If there is any delay in the downstream libraries updating, it delays.

    I see this as the original motivating factor behind this, and I know the pain. However, as someone already mentioned, I think this is the best use-case for AI tools, which greatly diminish, if not outright remove all the friction associated with this process. Thus I'd question whether DataFusion even needs to commit to less frequent major / breaking releases in the first place (in this day and age)?

    Arguably, AI tools might not be evenly distributed yet, and human review can still be a point of friction in the upgrade process, so the premise (less frequent major / breaking releases) might still be justifiable. In that case I'd favor option 1 rather than LTS branches, as I fear LTS branches shift some of the maintenance burden to consumer crates (delta-rs/iceberg-rust/etc.), which might not have the resources/discipline to properly handle it, and so it would end up being a waste of effort on the DF community part.

    To be concrete, consider the following options the maintainers of the DF consumer crates have:

    1. (status quo) ignore the DF LTS versions, and just keep updating the main branch with the latest DF version when released
      • LTS effort wasted
    2. (wrong approach) start pinning the main branch to DF LTS versions
      • defer the upgrade effort to later (when it's much bigger)
      • doesn't really resolve the original problem, since it now blocks the user from picking up newer DF versions since some crates are stuck on LTS
    3. ("right" approach) maintain separate LTS branches of their own, that track DF LTS branches
      • more maintenance work than 1 or 2, they not only have to manage these branches/releases, but also likely do some backporting of their own fixes/features from their main branch
      • what about their main branch now?
        • If it just tracks DF LTS, then there's no point in a separate LTS branch to begin with, and we're back to case 2
        • but then again if it doesn't then it may as well track the newest DF version, and we're back to case 1 (thus defeating the purpose of a separate LTS branch and rendering the DF LTS effort wasted again)
      • defer the upgrade effort (of the LTS branch) to later (when it's much bigger)

    None of these options actually benefit the end user and the crate maintainers simultaneously, so I'm wondering what would be the canonical way of using DF LTS versions then?

  25. alamb commented on Sep 19, 2026

    @alamb
    ContributorAuthor

    Thank you @gruuya -- the point out deferring upgrade effort to later when it is bigger is a good one.

    I terms of more work for downstream crates, I am not sure it is any easier / harder on them as today they need to chase multiple DataFusion versions potentially anyways

    I would actually be happy with Option 1 as well personally (do less frequent breaking API releases) but I think that shifts the burden to new feature development. That might be the right shift for the project at this point in its life, but I am not sure everyone agrees with that assesment

  26. comphead commented on Sep 19, 2026

    @comphead
    Contributor

    For example, Apache Spark follows a SemVer-inspired versioning strategy, x.y.z, with some project-specific deviations:

    • x — Major release: roughly once a year. This is where breaking changes, API removals/deprecations, dependency upgrades, and other changes incompatible with the previous major line can happen. It is the point where significant migrations may be required.
    • y — Feature release: historically quarterly. These releases add new features, performance improvements, API additions, and bug fixes, while maintaining compatibility within the major line. Public APIs can be added, but existing public APIs are not changed or removed.
    • z — Maintenance release: released as needed for critical bug fixes and security/correctness fixes. These are intended to be patch-level, compatible updates.

    Spark's current release policy is evolving toward an annual major + quarterly feature-release cadence. The last feature release of each major line is designated as LTS and receives 18 months of maintenance. Spark 3.5.x is a special extended-LTS case, with that period ending in November 2027.

    That said, API stability also depends heavily on the project's maturity and evolution velocity. Spark is a mature project, whereas DataFusion is still evolving rapidly, so applying the same compatibility guarantees today may not be practical.

    But I agree that, eventually, DataFusion will probably reach a point where adopting a similar versioning and release strategy makes sense.

  27. gruuya commented on Sep 22, 2026

    @gruuya
    Contributor

    In terms of more work for downstream crates, I am not sure it is any easier / harder on them as today they need to chase multiple DataFusion versions potentially anyways

    Also worth hearing out from some maintainers here (re: option 2, i.e. LTS releases/branches), @JanKaul, @rtyler, @ion-elgreco and @gabotechs come to mind.

  28. gabotechs commented on Sep 22, 2026

    @gabotechs
    Contributor

    Giving my two cents: I can imagine how there's a subset of people that are fine with long-term LTS branches, because their system might not need a heavy use of new additions to the public API, their DataFusion-based systems might be on maintenance mode, or something similar.

    However, the public API of DataFusion is wide, and all the breaking changes introduced to it in the last year seem legitimate and valuable, something that "heavy" DataFusion users will likely want to benefit sooner rather than later. Even if I see the value in LTS releases, I'd be fine with just not having them if that allows the project to move measurably faster.

    My impression is that DataFusion's public API is not mature enough to easily afford LTS releases without a big impact on development speed, which is something that could be steered with some discipline in public API design, and revisited in the future. If breaking changes are properly documented in an upgrade guide (which is the case today), typically just throwing an LLM to it easily one-shots a full upgrade, so (at least IME) frequent breaking changes are not that bad.

    That being said, I don't have a strong personal opinion on neither of the options.

  29. rtyler commented on Sep 22, 2026

    @rtyler
    Contributor

    In terms of more work for downstream crates, I am not sure it is any easier / harder on them as today they need to chase multiple DataFusion versions potentially anyways

    Also worth hearing out from some maintainers here (re: option 2, i.e. LTS releases/branches), @JanKaul, @rtyler, @ion-elgreco and @gabotechs come to mind.

    crawls out of the sewers

    I have been summoned! From the delta-io/delta-rs perspective Datafusion major releases are tedious but not painfully so, it is usually a case of "who moved my cheese" trying to figure out which APIs moved around and what they mean. ("This was a PlanConfiguration now it's an ConfiguredExecPlan? okie doke")

    The arrow and object_store breaking changes have been excruciating because we basically have to align the planets between delta-io/delta-kernel-rs, which must necessarily support multiple arrow versions through feature flags, all the way up the stack through datafusion, delta-rs, and then some of the extended ecosystem as well (e.g. datafusion-ffi / python). @comphead's point about the Spark release cycle is the de facto world we have to live in because these planets align typically only once or twice a year for major release versions.

    The thing about API breaking changes is that for most users, myself included, the trtadeoff is better performance or capability, most of us will jump through as many hoops as you present in order to get those new capabilities 😄

    I appreciate the ping on the thread and the consideration here 🫡

  30. alamb commented on Sep 22, 2026

    @alamb
    ContributorAuthor

    Maybe we could try the API stability thing for a release and see how bad it would be. I am not sure we'll know if we try it

    For example, we could say that for the 4 week-6week release cycle after DataFusion 56 is released, we won't commit breaking changes to main and then release 56.1.x directly from main 🤔

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions