Repository navigation
JpegDecoder: post-process baseline spectral data per MCU-row #1597
Description
Activity
I can definitely get behind this. As I recall looking at other libraries they were working on a per MCU process but I think per MCU row is fine.
It's a shame this optimization is only limited to sequential jpegs but is definitely worth the effort.
- added and removed
on Apr 14, 2021 I hope to be able to have a look at this within ~3 weeks, unless someone else wants to take it earlier.
@br3aker is this something you'd be interested in? You've been doing amazing work in the encoder!
@JimBobSquarePants I can definately get behind this after PR with that deBruijn table I was talking about at memory allocator PR.
Reacted by James Jackson-South, Jason Nelson and Anton FirszovDid a little study to understand decoder architecture. Long story short: code explosion due to generic
TPixeltype propagation. Current workflow for jpeg encoder:Decode ... ParseStream<TPixel> ... Parse Start of Frame // resulting image size is known - allocate Buffer2D instead of in post process step Parse Start of Scan(s) // that's where the problem lies - there's no way convert spectral data to TPixel without generics ... ... PostProcessIntoImage<TPixel> // spectral -> YCbCr -> Rgba -> TPixelConverting spectral data to YCbCr and then to Rgba is a piece of cake as we know for sure that supported jpeg contains spectral data which would yield YCbCr colorspace values which should be converted to Rgba for future colorspace conversion - no generics needed. Rgba -> TPixel is done via
PixelOperations<TPixel>.Instance.FromVector4Destructive(...)call which is:- Impossible to get without making entire call stack generic
- I guess any
PixelOperations<TPixel>.Instancevirtual calls are devirtualized, at least in jpeg encoder it's not a bottleneck
There's no need to convert full mcu row as PostProcessor converts them piece by piece so it's unlikely to bring any performance benefits. Moreover, 4:2:0 needs to process two rows at the same time, so 2 full rows of
Block8x8must be allocated. I think that per block decoding directly to the Image is the best approach here.While machine code size can bloat at runtime for each pixel type decoded from jpeg, I don't think that majory of users would use more than 1-3 pixel types they want images to decode to.
@JimBobSquarePants @antonfirsov I might overlooked something but I'm almost confident that this is the only way, can you elaborate on the final decision?
@br3aker I think we can solve this with a double-dispatch trick by implementing
IImageVisitorsomewhere insideJpegImagePostProcessor. After initializing it with anImageinstance, you can define a non- generic method to consume the spectral data, and callimage.Accept(visitor)to delegate the implementation to a generic method. You can then call this non-generic method fromHuffmanScanDecoder. In the end, the genericIImageVistor.Visit<T>implementation will be very similar to whatPostProcess<T>(imageFrame)does.The recommendation to go MCU row by MCU-row is mostly to avoid the overhead of virtual calls. Note that virtual methods on
PixelOperations<TPixel>can not be devirtualized, calling them with block granularity would be inefficient (same for the color converters).Reacted by DmitryYep, was a bit delusional about the power of the JIT :D. Thanks for pointing that out.
I though as TPixel is a struct, jit would compile IL to an exclusive implementation for exact TPixel type. Completely forgot that
PixelOperations<TPixel>is the base class, not an actual implementation.Won't be able to work for a couple of days but will definitely work on this, thanks for the double dispatch advice!
I think that per block decoding directly to the Image is the best approach here.
@br3aker Fairly certain libjpeg turbo etc do it one block at a time for baseline. I don't know if there's an optimization we can do for progressive.@JimBobSquarePants I meant one block at a time but for a bulk of mcus at the same time so it would eliminate virtual call overhead:
foreach stride in image: foreach mcu in stride: rgbaMcu = mcu.FromSpectral().ToRgba(); SomeMagicClass.ConvertToImage(rgbaMcu); // virtual call convertion per each 8x8 mcu in imagevs.
rgbaStride = allocator.Alloc(mcusPerStride); foreach stride in image: foreach mcu in stride: rgbaMcuStride[i] = mcu.FromSpectral().ToRgba(); SomeMagicClass.ConvertToImage(rgbaMcuStride); // virtual call convertion per each mcu stride in imageThe only problem here is memory allocation for mcu stride, especially for 420 subsampling as it proccesses more mcus per decoding unit (4 luma + 1 chroma) per actual deconding to spectral data. Maybe it's better to process in more granular bulks depending on some allocator size like it was 2MB last discussion but it doesn't matter that much atm so we can discuss it later when at least mvp implementation is ready.
Encoder actually has the same problem, it calls
PixelOperations<TPixel>.Instance.Convert(,,,)for each mcu in image, I'll benchmark it for bulk conversion after decoder stuff, maybe it's still possible to squeez some more performance out of it.12 remaining items
@antonfirsov encountered a little dilemma:
For baseline jpeg files proposed conversion is straightforward as it's known when spectral stride for each component is ready.
For progressive it's s bit tricky as progressive can define separate scans for each component soif(spectralEnd == 63) /* Convert spectral to color here */won't work. There are two solutions:public Image<TPixel> Decode<TPixel>(BufferedReadStream stream, CancellationToken cancellationToken) where TPixel : unmanaged, IPixel<TPixel> { // this is still WIP but final variant would look somewhat like this var specificConverter = new SpectralToImageConverter<TPixel>(this.Configuration); this.spectralConverter = specificConverter; this.ParseStream(stream, cancellationToken: cancellationToken); this.InitExifProfile(); this.InitIccProfile(); this.InitIptcProfile(); this.InitDerivedMetadataProperties(); // this looks out of place to be honest if (/* This jpeg is progressive */) { specificConverter.ConvertFullScan(); } return new Image<TPixel>(this.Configuration, this.Metadata, new[] { specificConverter.ImageFrame }); }
Another solution is to comit spectral data to the converter before returning from
ParseStream()method.While both solutions look a bit 'ugly' it's the most performant way of checking
are all scans done?. What do you think?P.S.
Yes,Image<TPixel>creation looks ugly, I will open a separate PR for new ctor from single frame soon.@br3aker I like the plan with the pseudo
Decode<TPixel>method. I don't think there is a way to avoid this complexity (or "ugliness"), since it comes from the jpeg standard itself.@antonfirsov I will hide if-check in property getter for visual clarity then. Thanks for the responce!
A little update on this:
I redid a lot of code and screwed something up and I couldn't find out why in a couple of hours + got a new idea which should be a little more understandable.
First of all, resulting image whould be constructed from
Buffer2D<TPixel>buffer. This would introduce a newImage<TPixel>constructor but it would be less intrusive. @antonfirsov not sure what to do in with #1680, implement pixel buffer ctor there or within this issue PR?Second, I've decided to implement this as an enumerable collection of spectral strides:
// There would actually be some wrapping class for strides // so it would store all components' spectral strides in a single object foreach(Buffer2D<Block8x8> spectralStride in scanDecoder) { // spectral -> vector4 Buffer2D<Vector4> colorBuffer = ConvertFromSpectalToVector4(spectralStride); // vector4 -> TPixel PixelBuffer[i] = ConvertFromVector4<TPixel>(colorBuffer); }
For every decoding mode it won't change anything (except there's won't be extra virtual call for each stride conversion) but for baseline dct it would allow to deffer stream parsing stride by stride.
Reacted by James Jackson-South@antonfirsov not sure what to do in with #1680, implement pixel buffer ctor there or within this issue PR?
If you have a working PR, that would be a good trigger to push things into a decision, otherwise it's just endless
bike shaddingAPI discussions 😄Second, I've decided to implement this as an enumerable collection of spectral strides.
foreach(Buffer2D<Block8x8> spectralStride in scanDecoder)I wonder how does this work with
ProcessStartOfScan? Does SOS kick of the code that iterates through the enumerable stuff, reading the stream further internally? What if there are more SOS-s?Reacted by James Jackson-SouthWhat if there are more SOS-s?
We definitely need to cater for that since that's how progressive jpegs work.
I would open a draft PR where we can discuss the actual implementation.
I wonder how does this work with
ProcessStartOfScan? Does SOS kick of the code that iterates through the enumerable stuff, reading the stream further internally? What if there are more SOS-s?That's not that hard to determine actually. Multiple SOS markers can exist only in:
- progressive jpegs - can be checked via SOF2 marker existence, current internal API already has it
- non-interleaved jpegs - can be checked if given scan component count is not equal to 'global' component count defined in SOF segment
In other words:
if (!this.Frame.IsProgressive && this.Frame.ComponentCount == scanComponentCount) { // this SOS must be the only one, any extra is an error and can be checked after spectral decoding this.scanDecoder.Baseline = true; // we can return true to signal that we are ready for spectral conversion return true; } // decodes current partial scan to the pre-allocated spectral buffer this.scanDecoder.DecodeScan(); // for more consistent behaviour we can actually evaluate if multi-sos jpeg is done // via spectralEnd == 63 for progressive jpegs // via processedScans == this.Frame.ComponentCount for non-inteleaved baseline jpegs return lastScanCondition;
There's a problem if given jpeg has anything after SOS except EOI and we can actually check even that - we can call ParseStream() one more time.
Nevermind, current architecture & code is not in a good shape for my plan, I'll try to work on it later. Priority right now:
- Working PR fixing this issue with least code change possible
- PR for refactoring (a lot of decoupling needed for decoder core & scan decoder) and micro-optimizations (there are some good places to cut off couple of ms)
- (possible) PR for Enumerable idea
Sorry for this rapid change of ideas & messages, yesterday's discard of an almost working code knocked me hard. Will try to work out an implementation in a couple of days.
Reacted by James Jackson-South and Anton Firszovyesterday's discard of an almost working code knocked me hard
No need to apologize and I feel you mate. At some point I need to replay months of optimization code I wrote for a zlib stream implementation because at some point I broke it but have no idea when. 😞
Really looking forward to seeing what you come up with!
Reacted by Dmitry@antonfirsov @JimBobSquarePants sorry for bothering but I have a little problem
PR is actually ready with almost all tests passing. The only problem is these tests for baseline jpegs:
These tests check if spectral data is equal to libjpeg spectral data of the given image. This approach is impossible as spectral data is discarded deep inside scan decoding process. Progressive and multi-scan baselines can be tested simply because they use the same technique as before the PR.
Question: Do we even need to test spectral data? We can compare final colors which would be invalid if spectral is invalid. Yes, it's a couple layers 'higher' but right now it's impossible to test. Only if
Enumerableapproach will be implemented - spectral strides would be given out by virtual method so it would be possible to inject testing code.P.S.
jpeg400jfif.jpgmust fail - it's a known bug, don't mind it.We can compare final colors which would be invalid if spectral is invalid.
The good thing about verifying spectral data is that the spectral intermediate result is exact, while for color conversion small deviations are allowed. It's important to be able to catch cases when a small difference is a result of a Huffman-decoding bug and not a floating point inaccuracy. We faced and fixed such issues while finalizing #274, and while I don't expect a refactor of that volume anytime soon, the tests can be still handy for Huffman decoding optimizations, so I prefer to keep them in long term.
On the other hand, this should not block progress, so my recommendation is to temporarily disable them, and re-enable when the Enumerable refactor is done. @JimBobSquarePants agreed?
@antonfirsov skipping tests now seems like an easy plan but you know...
I will slightly alter current PR architecture to enable testing without major changes so it won't rely on 'somewhat possible enumerable implementation in the future'.
Reacted by Anton FirszovI will slightly alter current PR architecture to enable testing without major changes so it won't rely on 'somewhat possible enumerable implementation in the future'
Not sure I follow what you mean here? Do you mean that this is not an issue now? I'm happy to temporarily disable baseline tests for now simply to see a difference.
Not sure I follow what you mean here? Do you mean that this is not an issue now? I'm happy to temporarily disable baseline tests for now simply to see a difference.
It is a problem because I wanted to change as little code as possible for this PR so you guys won't spend too much time reviewing. These tests fix would result in a bigger change than necessary for PR to work. I will disable them and push a draft PR then.
Reacted by James Jackson-South and Anton Firszov

Currently, Huffmann decoding (done by
HuffmanScanDecoder) is strictly separated from postprocessing/color conversion (done byJpegImagePostProcessor) for simplicity. This means thatJpegComponent.SpectralBlocksare allocated upfront for the whole image.I did a second round of memory profiling using
SimpleGcMemoryAllocatorto get rid of pooling for more auditable results. This shows thatSpectralBlocksare responsible for the majority of our memory allocation:This can be eliminated with some non-trivial, but still limited refactoring:
JpegDecoderCoreandHuffmanScanDecoderneeds a mode whereJpegComponent.SpectralBlocksis interpreted as a sliding window of blocks instead of full set of decoded spectral blocksHuffmanScanDecodercan then push the rows of the sliding window, directly calling an instance ofJpegComponentPostprocessorin the end of it's MCU-row decoding loop