Repository navigation
[DISCUSS] DataFusion less frequent major / breaking releases (ease using multiple third-party extensions (like delta, or iceberg) ) #16622
Description
Activity
- changed the title
[-][DISCUSS] DataFusion minor releases (ease using multiple third-party extensions (like delta, or iceberg) )[/-][+][DISCUSS] DataFusion less frequent major / breaking releases (ease using multiple third-party extensions (like delta, or iceberg) )[/+]on Jun 30, 2025 I would like downstream libraries to have more time and schedule flexibility when upgrading DataFusion and other dependent crates, so that it is easier to construct a system from different components
That's a noble goal.
At the same time we should avoid policies that slow down DataFusion development.
Perhaps part of a solution could be to invert dependencies between the project. For example, Iceberg-DataFusion and Delta-DataFusion integrations could live in this repo.
Perhaps part of a solution could be to invert dependencies between the project. For example, Iceberg-DataFusion and Delta-DataFusion integrations could live in this repo.
That would make running tests in the DataFusion repo harder, perhaps (certainly there would be more of them).
Also, while it could work for icerberg and delta-rs potentially, I am not sure we could plausibly bring all the integrations into the same repo (though I will admit it seems to work for trino 🤔 )
I am not opposed necessairly, just trying to think of the trade offs
Another approach maybe could be to break the
datafusionrepo into separate repos with less stuff that can move slower and keep as much as possible in other repos that can move faster 🤔IMHO, option 1 may be better, it is not as frequent as current release cycle but still frequent enough to take breaking changes in relatively small and digestible chunks.
Not sure if option 2 would force downstream project to have same branching strategy. Otherwise it would bring "big update" every time LTS was released.
Reacted by Andrew LambAnd to be clear, in my mind option 1 still has us doing monthly releases, we just restrict which releases have breaking API changes
Reacted by Marko MilenkovićI tend to favor option 1. I am currently in this same position where each version of DF requires updates to Lance and datafusion-python. In the last update, within a week of getting my repo upgraded to DF47, DF48 was released.
The part that is unclear to me, but maybe arrow-rs has already solved, is how we go about going PR reviews on work that does introduce breaking changes. Consider the timeline where a feature is developed and ready for review two weeks after a major release. That feature contains breaking API changes. Do we then require that feature to sit in a long living PR for another 6-8 weeks until we merge it to main? Those can be extremely difficult to maintain.
If we look at the problem through the lens of project consumers, the expectation to never break anything is natural. If we look at the problem of feature implementers, the exception to be able to introduce changes, even those that are unfortunately breaking, is natural.
arrow-rs has 2.6x less code and conceptually much simpler API surface. DataFusion is a library for building query engines and data processing tools. It is in active development and is not yet in "mostly complete, mostly maintenance" stage, so internal API changes are often unavoidable. DataFusion being a library, it has enormous API surface, so most of "internal API changes" are actually potentially breaking changes for downstream consumers.
Being able to do such changes once a quarter (4 times a year), looks too rare to me.
Also, while it could work for icerberg and delta-rs potentially, I am not sure we could plausibly bring all the integrations into the same repo
Agreed. At the same time, this issue is not on context of any integrations. It's in a context of real-life problems we (some of us) have today. What are these problems, and what integrations they pertain?
This is a pretty sweet idea from @jonmmease about making upgrades easier (use LLM agents): #13648 (comment)
(it is somewhat orthogonal to the other goals of this ticket, but I thought it was so good I wanted to share. Maybe like all great ideas it seems obvious in retrospect)
I favor option 2 (as already pointed out in the original issue). I don't think that we should artificially slow down development against the
mainbranch.We already have a documented process for backporting PRs to release branches (e.g.
branch-48) and creating future 48.x.x releases without breaking changes. We always release from the release branches, not themainbranch. This is the same model that Apache Spark uses. We just haven't been proactive in backporting PRs to the release branches most of the time.The idea of holding off on merging PRs to
mainsounds like a nightmare to me. In Comet, we often temporarily point to a fixed revision of DataFusion'smainbranch so that we can gradually upgrade to the next version, testing all recently merged PRs. If breaking change PRs are all sitting in a queue until just before the next major DataFusion release, then I'm not sure how we can gradually upgrade.The responsibility for creating the PRs to backport fixes to the release branch should fall to the downstream users who are waiting on those fixes. There is additional work for DataFusion maintainers in terms of PR reviews though.
Reacted by Piotr Findeisen, Oleks V, Andrew Lamb and Marko MilenkovićMore inclining to version 2 with slight modifications. Similar to Rustc flow https://web.mit.edu/rust-lang_v1.25/arch/amd64_ubuntu1404/share/doc/rust/html/book/second-edition/ch01-03-how-rust-is-made-and-nightly-rust.html
so it would be a
mainbranch representing latest update(keep in mind main can be broken by bad commits, performance issue, etc)a
nightlyrelease branch where DF releases artifacts per night or per week, this branch should be protected by extended tests, performance tests, etcSo for Version 2 (a LTS branch) do we have any proposal for the cadence of releases?
Like would we do 2 releases each month now?
- LTS release
- Major release from main
Or would we perhaps keep one release a month and release on alternate months 🤔
if I may, duckdb process seems a bit reasonable, release a major version with breaking change every three months (1.1, 1.2 , 1.3 etc), release a bug fix every month, 1.3.1 , 1.3.2 then next version will be 1.4, and they have nightly build for both the main branch and the stable branch
Reacted by Andrew Lambif I may, duckdb process seems a bit reasonable, release a major version with breaking change every three months (1.1, 1.2 , 1.3 etc), release a bug fix every month, 1.3.1 , 1.3.2 then next version will be 1.4, and they have nightly build for both the main branch and the stable branch
This is basically the same cadence that we use in arrow-rs as well. The only potential issue I see with doing something like that for DataFusion is accumulating 3 months of breaking changes might be a lot 🤔
Reacted by Mimoune14 remaining items
Can we consider "breaking change per quarter budget", it would raise a bar for merging braking changes,
Wouldn't this directly interfere with the comment above:
I believe major releases every six to eight weeks helped datafusion growth
Meaning how much of the growth are we leaving on table because we ran out of breaking change budget? Also what would be the criteria for accepting/rejecting a breaking changes and how do we track them.
We do not know what the percentage of breaking changes were made as there were no other options and what percentage ware done as that was the easy path forward, having some kind of budget could create a pushback on later.
Criteria for accepting breaking change was deliberately left out for discussion, it could be number of pmc or mainintainer endorsement or similar
- I would like to keep releasing major versions every 6-8 weeks, so that downstream projects such as Comet can keep moving forward without the need to start forking. I am +1 for maintaining LTS releases for downstream consumers who need stability. I would like Comet to also maintain LTS releases at some point.…On Sat, Sep 12, 2026 at 9:33 AM Marko Milenković ***@***.***> wrote: *milenkovicm* left a comment (apache/datafusion#16622) <#16622 (comment)> Can we consider "breaking change per quarter budget", it would raise a bar for merging braking changes, Wouldn't this directly interfere with the comment above: I believe major releases every six to eight weeks helped datafusion growth Meaning how much of the growth are we leaving on table because we ran out of breaking change budget? Also what would be the criteria for accepting/rejecting a breaking changes and how do we track them. We do not know what the percentage of breaking changes were made as there were no other options and what percentage ware done as that was the easy path forward, having some kind of budget could create a pushback on later. Criteria for accepting breaking change was deliberately left out for discussion, it could be number of pmc or mainintainer endorsement or similar — Reply to this email directly, view it on GitHub <#16622?email_source=notifications&email_token=AAHEBRCHPM2FVZ6Q5NK7E7D5OVUFZA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKNRUGY4DKNBXGEY2M4TFMFZW63VHNVSW45DJN5XKKZLWMVXHJLDGN5XXIZLSL5RWY2LDNM#issuecomment-5646854711>, or unsubscribe <https://github.com/notifications/unsubscribe-auth/AAHEBRHEETZHZMYNMCOVYRL5OVUFZAVCNFSNUABFKJSXA33TNF2G64TZHMZTKOBZGE3TGMJYHNEXG43VMU5TGMJYHAYDONRQHE32C5QC> . You are receiving this because you were mentioned.Message ID: <apache/datafusion/issues/16622/5646854711 ***@***.***>Reacted by Kumar Ujjawal
In some ways our organization is already doing this - I typically update our internal datafusion version every 3rd or 4th release as I just don't have the bandwidth to test every single release. I'd love a LTS version for that reason if the release cycle for that was every 6-12 months.
I'd love a LTS version for that reason if the release cycle for that was every 6-12 months.
I feel the same. Even though there would be a lot of work involved for the people involved in the release process. It would be a good step to have TLS.
Reacted by Andrew LambSomewhat related to this, I think it would be nice if a merged PR had a way to quickly tell which version it applies to (or will apply). Sometimes in the past I had to check the changelog to confirm if a PR got in the current release or not. We could use milestones (like Rust or Kubernetes), or regular labels (like Helm):
This should be easy to do with actions that run on merge.
If then DataFusion supports an LTS version, we could add a label there as well so it is easy to tell that it was backported.
@alamb High-level, an LTS in some form could definitely make sense. Coming from the more traditional database world, the fact that we have thus far gotten away with shipping API-breaking releases every 6-8 weeks is a little surprising 😅
Indeed!
- Six months is a pretty short support duration for an LTS release. Is that long enough that it will sufficiently ease downstream pain?
I think it would ease downstream pain compared to 1.5 months we have now 😅 -- but for sure the longer we supported releases the easier it would be for downstream projects to sync.
- What should the criteria be for backporting changes to the LTS branch?
I would initially propose a pretty lenient policy (basically anything that doesn't have API changes and is not deemed "too risky" by maintainers). For example, we probably wouldn't backport refactors or major new features like
AS OFJoins. However, if we find we are introducing too many regressions we could tighten up the backport policy.But maybe coding agents make it more feasible to do aggressive backporting?
This is my thesis and I think we could try it for a while
- Should we consider shipping DataFusion releases less often?
This is also a possibility (though I would personally suggest we go to the major/minor release model described in this issue's description as it has worked well for arrow). I think it would work just as well, but would shift the burden as it would restrict development on
mainI typically update our internal datafusion version every 3rd or 4th release as I just don't have the bandwidth to test every single release. I'd love a LTS version for that reason if the release cycle for that was every 6-12 months.
@Omega359 we have basically done this at InfluxData too (not deliberately, but this is what has sort of happened organically)
I think it would be nice if a merged PR had a way to quickly tell which version it applies to (or will apply). Sometimes in the past I had to check the changelog to confirm if a PR got in the current release or not. We could use milestones (like Rust or Kubernetes), or regular labels (like Helm):
@nuno-faria I agree this would be super helpful and is not clear from PRs at the moment. Maybe it is worth starting another issue on this topic 🤔
Reacted by Nuno FariaCan we consider "breaking change per quarter budget", it would raise a bar for merging braking changes, yet it would allow controlled number of braking change to get in.
@milenkovicm this is a very interesting idea and one I think we should pursue in parallel -- maybe we could start with a "really don't break the API unless needed" type wording for PRs (aka really emphasize using deprecation rather than changing signatures for example)
I noticed the idea of DF LTS versions being discussed in the context of the DF 56.0.0 release. As someone who's gone through many many DF updates (we haven't missed a single one since DF 9.0.0!), I'd like to share my comments.
This means for an application that relies on multiple downstream packages must wait until ALL of them have upgraded to the new version in order to upgrade DataFusion. If there is any delay in the downstream libraries updating, it delays.
I see this as the original motivating factor behind this, and I know the pain. However, as someone already mentioned, I think this is the best use-case for AI tools, which greatly diminish, if not outright remove all the friction associated with this process. Thus I'd question whether DataFusion even needs to commit to
less frequent major / breaking releasesin the first place (in this day and age)?Arguably, AI tools might not be evenly distributed yet, and human review can still be a point of friction in the upgrade process, so the premise (
less frequent major / breaking releases) might still be justifiable. In that case I'd favor option 1 rather than LTS branches, as I fear LTS branches shift some of the maintenance burden to consumer crates (delta-rs/iceberg-rust/etc.), which might not have the resources/discipline to properly handle it, and so it would end up being a waste of effort on the DF community part.To be concrete, consider the following options the maintainers of the DF consumer crates have:
- (status quo) ignore the DF LTS versions, and just keep updating the main branch with the latest DF version when released
- LTS effort wasted
- (wrong approach) start pinning the main branch to DF LTS versions
- defer the upgrade effort to later (when it's much bigger)
- doesn't really resolve the original problem, since it now blocks the user from picking up newer DF versions since some crates are stuck on LTS
- ("right" approach) maintain separate LTS branches of their own, that track DF LTS branches
- more maintenance work than 1 or 2, they not only have to manage these branches/releases, but also likely do some backporting of their own fixes/features from their main branch
- what about their main branch now?
- If it just tracks DF LTS, then there's no point in a separate LTS branch to begin with, and we're back to case 2
- but then again if it doesn't then it may as well track the newest DF version, and we're back to case 1 (thus defeating the purpose of a separate LTS branch and rendering the DF LTS effort wasted again)
- defer the upgrade effort (of the LTS branch) to later (when it's much bigger)
None of these options actually benefit the end user and the crate maintainers simultaneously, so I'm wondering what would be the canonical way of using DF LTS versions then?
Reacted by Jay Zhan and Andrew Lamb- (status quo) ignore the DF LTS versions, and just keep updating the main branch with the latest DF version when released
Thank you @gruuya -- the point out deferring upgrade effort to later when it is bigger is a good one.
I terms of more work for downstream crates, I am not sure it is any easier / harder on them as today they need to chase multiple DataFusion versions potentially anyways
I would actually be happy with Option 1 as well personally (do less frequent breaking API releases) but I think that shifts the burden to new feature development. That might be the right shift for the project at this point in its life, but I am not sure everyone agrees with that assesment
Reacted by Marko GrujicFor example, Apache Spark follows a SemVer-inspired versioning strategy,
x.y.z, with some project-specific deviations:x— Major release: roughly once a year. This is where breaking changes, API removals/deprecations, dependency upgrades, and other changes incompatible with the previous major line can happen. It is the point where significant migrations may be required.y— Feature release: historically quarterly. These releases add new features, performance improvements, API additions, and bug fixes, while maintaining compatibility within the major line. Public APIs can be added, but existing public APIs are not changed or removed.z— Maintenance release: released as needed for critical bug fixes and security/correctness fixes. These are intended to be patch-level, compatible updates.
Spark's current release policy is evolving toward an annual major + quarterly feature-release cadence. The last feature release of each major line is designated as LTS and receives 18 months of maintenance. Spark 3.5.x is a special extended-LTS case, with that period ending in November 2027.
That said, API stability also depends heavily on the project's maturity and evolution velocity. Spark is a mature project, whereas DataFusion is still evolving rapidly, so applying the same compatibility guarantees today may not be practical.
But I agree that, eventually, DataFusion will probably reach a point where adopting a similar versioning and release strategy makes sense.
Reacted by Jay Zhan, Marko Grujic and Andrew LambIn terms of more work for downstream crates, I am not sure it is any easier / harder on them as today they need to chase multiple DataFusion versions potentially anyways
Also worth hearing out from some maintainers here (re: option 2, i.e. LTS releases/branches), @JanKaul, @rtyler, @ion-elgreco and @gabotechs come to mind.
Giving my two cents: I can imagine how there's a subset of people that are fine with long-term LTS branches, because their system might not need a heavy use of new additions to the public API, their DataFusion-based systems might be on maintenance mode, or something similar.
However, the public API of DataFusion is wide, and all the breaking changes introduced to it in the last year seem legitimate and valuable, something that "heavy" DataFusion users will likely want to benefit sooner rather than later. Even if I see the value in LTS releases, I'd be fine with just not having them if that allows the project to move measurably faster.
My impression is that DataFusion's public API is not mature enough to easily afford LTS releases without a big impact on development speed, which is something that could be steered with some discipline in public API design, and revisited in the future. If breaking changes are properly documented in an upgrade guide (which is the case today), typically just throwing an LLM to it easily one-shots a full upgrade, so (at least IME) frequent breaking changes are not that bad.
That being said, I don't have a strong personal opinion on neither of the options.
In terms of more work for downstream crates, I am not sure it is any easier / harder on them as today they need to chase multiple DataFusion versions potentially anyways
Also worth hearing out from some maintainers here (re: option 2, i.e. LTS releases/branches), @JanKaul, @rtyler, @ion-elgreco and @gabotechs come to mind.
crawls out of the sewers
I have been summoned! From the delta-io/delta-rs perspective Datafusion major releases are tedious but not painfully so, it is usually a case of "who moved my cheese" trying to figure out which APIs moved around and what they mean. ("This was a PlanConfiguration now it's an ConfiguredExecPlan? okie doke")
The
arrowandobject_storebreaking changes have been excruciating because we basically have to align the planets between delta-io/delta-kernel-rs, which must necessarily support multiple arrow versions through feature flags, all the way up the stack through datafusion, delta-rs, and then some of the extended ecosystem as well (e.g. datafusion-ffi / python). @comphead's point about the Spark release cycle is the de facto world we have to live in because these planets align typically only once or twice a year for major release versions.The thing about API breaking changes is that for most users, myself included, the trtadeoff is better performance or capability, most of us will jump through as many hoops as you present in order to get those new capabilities 😄
I appreciate the ping on the thread and the consideration here 🫡
Reacted by Ion KoutsourisReacted by Andrew LambMaybe we could try the API stability thing for a release and see how bad it would be. I am not sure we'll know if we try it
For example, we could say that for the 4 week-6week release cycle after DataFusion 56 is released, we won't commit breaking changes to
mainand then release 56.1.x directly from main 🤔Reacted by R. Tyler Croy
Is your feature request related to a problem or challenge?
One of the dreams of the composable data ecosystem is to quickly assemble a system from various components (DataFusion, data formats
DataFusion still releases once a month, which allows code to quickly flow but also causes at least 2 challenges:
Third party extensions like delta-rs and iceberg provide
TableProvidersfor DataFusion, which is really nice. However, to use those packages the versions of DataFusion must match exactly.This means for an application that relies on multiple downstream packages must wait until ALL of them have upgraded to the new version in order to upgrade DataFusion. If there is any delay in the downstream libraries updating, it delays.
For example, an application that wants to use delta-rs, iceberg, and the
table-providerscrate, there is a race after each upgrade of DataFusionLet's take a release timeline for
XreleasedXXDescribe the solution you'd like
I would like downstream libraries to have more time and schedule flexibility when upgrading DataFusion and other dependent crates, so that it is easier to construct a system from different components
Describe alternatives you've considered
Option 1: Switch to major/minor release cadence
We could follow the model of arrow-rs which does releases monthly, but breaking releases only quarterly. Here is how it works in arrow-rs: https://github.com/apache/arrow-rs?tab=readme-ov-file#release-versioning-and-schedule
This would mean continuing to release every month, but only allowing breaking API changes every 3rd release (or some other cadence)
The major cost here is that maintainers and contributors would have to be diligent about not merging breaking API changes until a major release
This is possible to automate somewhat:
Option 2: LTS and feature branch
-Keep (at least) two branches going: LTS and main, as proposed by @andygrove in #5269
In this model we would likely backport changes to the LTS branch and make releases from there. The downside of this approach is that there is extra work to backport changes to LTS.
Option 3: Split API into Stable and Unstable parts
This is described by @thinkharderdev in
The idea is to extract the API into separate crates (e.g.
datafusion-catalog-api) that don't change as frequently and where the changes are more carefully controlledAdditional context
No response