Skip to content

JpegDecoder: post-process baseline spectral data per MCU-row #1597

Description

@antonfirsov

Currently, Huffmann decoding (done by HuffmanScanDecoder) is strictly separated from postprocessing/color conversion (done by JpegImagePostProcessor) for simplicity. This means that JpegComponent.SpectralBlocks are allocated upfront for the whole image.

I did a second round of memory profiling using SimpleGcMemoryAllocator to get rid of pooling for more auditable results. This shows that SpectralBlocks are responsible for the majority of our memory allocation:

image

This can be eliminated with some non-trivial, but still limited refactoring:

  • JpegDecoderCore and HuffmanScanDecoder needs a mode where JpegComponent.SpectralBlocks is interpreted as a sliding window of blocks instead of full set of decoded spectral blocks
  • HuffmanScanDecoder can then push the rows of the sliding window, directly calling an instance of JpegComponentPostprocessor in the end of it's MCU-row decoding loop

Activity

  1. JimBobSquarePants commented on Apr 14, 2021

    @JimBobSquarePants
    Member

    I can definitely get behind this. As I recall looking at other libraries they were working on a per MCU process but I think per MCU row is fine.

    It's a shame this optimization is only limited to sequential jpegs but is definitely worth the effort.

  2. added this to the Future milestone on Apr 14, 2021
  3. antonfirsov commented on Apr 16, 2021

    @antonfirsov
    MemberAuthor

    I hope to be able to have a look at this within ~3 weeks, unless someone else wants to take it earlier.

  4. JimBobSquarePants commented on Jun 14, 2021

    @JimBobSquarePants
    Member

    @br3aker is this something you'd be interested in? You've been doing amazing work in the encoder!

  5. br3aker commented on Jun 14, 2021

    @br3aker
    Contributor

    @JimBobSquarePants I can definately get behind this after PR with that deBruijn table I was talking about at memory allocator PR.

  6. br3aker commented on Jun 20, 2021

    @br3aker
    Contributor

    Did a little study to understand decoder architecture. Long story short: code explosion due to generic TPixel type propagation. Current workflow for jpeg encoder:

    Decode
        ...
        ParseStream<TPixel>
            ...
            Parse Start of Frame // resulting image size is known - allocate Buffer2D instead of in post process step
            Parse Start of Scan(s) // that's where the problem lies - there's no way convert spectral data to TPixel without generics
            ...
        ...
        PostProcessIntoImage<TPixel> // spectral -> YCbCr -> Rgba -> TPixel
    

    Converting spectral data to YCbCr and then to Rgba is a piece of cake as we know for sure that supported jpeg contains spectral data which would yield YCbCr colorspace values which should be converted to Rgba for future colorspace conversion - no generics needed. Rgba -> TPixel is done via PixelOperations<TPixel>.Instance.FromVector4Destructive(...) call which is:

    1. Impossible to get without making entire call stack generic
    2. I guess any PixelOperations<TPixel>.Instance virtual calls are devirtualized, at least in jpeg encoder it's not a bottleneck

    There's no need to convert full mcu row as PostProcessor converts them piece by piece so it's unlikely to bring any performance benefits. Moreover, 4:2:0 needs to process two rows at the same time, so 2 full rows of Block8x8 must be allocated. I think that per block decoding directly to the Image is the best approach here.

    While machine code size can bloat at runtime for each pixel type decoded from jpeg, I don't think that majory of users would use more than 1-3 pixel types they want images to decode to.

    @JimBobSquarePants @antonfirsov I might overlooked something but I'm almost confident that this is the only way, can you elaborate on the final decision?

  7. antonfirsov commented on Jun 20, 2021

    @antonfirsov
    MemberAuthor

    @br3aker I think we can solve this with a double-dispatch trick by implementing IImageVisitor somewhere inside JpegImagePostProcessor. After initializing it with an Image instance, you can define a non- generic method to consume the spectral data, and call image.Accept(visitor) to delegate the implementation to a generic method. You can then call this non-generic method from HuffmanScanDecoder. In the end, the generic IImageVistor.Visit<T> implementation will be very similar to what PostProcess<T>(imageFrame) does.

    The recommendation to go MCU row by MCU-row is mostly to avoid the overhead of virtual calls. Note that virtual methods on PixelOperations<TPixel> can not be devirtualized, calling them with block granularity would be inefficient (same for the color converters).

  8. br3aker commented on Jun 20, 2021

    @br3aker
    Contributor

    Yep, was a bit delusional about the power of the JIT :D. Thanks for pointing that out.

    I though as TPixel is a struct, jit would compile IL to an exclusive implementation for exact TPixel type. Completely forgot that PixelOperations<TPixel> is the base class, not an actual implementation.

    Won't be able to work for a couple of days but will definitely work on this, thanks for the double dispatch advice!

  9. JimBobSquarePants commented on Jun 21, 2021

    @JimBobSquarePants
    Member

    I think that per block decoding directly to the Image is the best approach here.
    @br3aker Fairly certain libjpeg turbo etc do it one block at a time for baseline. I don't know if there's an optimization we can do for progressive.

  10. br3aker commented on Jun 21, 2021

    @br3aker
    Contributor

    @JimBobSquarePants I meant one block at a time but for a bulk of mcus at the same time so it would eliminate virtual call overhead:

    foreach stride in image:
        foreach mcu in stride:
            rgbaMcu = mcu.FromSpectral().ToRgba();
            SomeMagicClass.ConvertToImage(rgbaMcu); // virtual call convertion per each 8x8 mcu in image
    

    vs.

    rgbaStride = allocator.Alloc(mcusPerStride);
    foreach stride in image:
        foreach mcu in stride:
            rgbaMcuStride[i] = mcu.FromSpectral().ToRgba();
        SomeMagicClass.ConvertToImage(rgbaMcuStride); // virtual call convertion per each mcu stride in image
    

    The only problem here is memory allocation for mcu stride, especially for 420 subsampling as it proccesses more mcus per decoding unit (4 luma + 1 chroma) per actual deconding to spectral data. Maybe it's better to process in more granular bulks depending on some allocator size like it was 2MB last discussion but it doesn't matter that much atm so we can discuss it later when at least mvp implementation is ready.

    Encoder actually has the same problem, it calls PixelOperations<TPixel>.Instance.Convert(,,,) for each mcu in image, I'll benchmark it for bulk conversion after decoder stuff, maybe it's still possible to squeez some more performance out of it.

  11. 12 remaining items

  12. br3aker commented on Jun 29, 2021

    @br3aker
    Contributor

    @antonfirsov encountered a little dilemma:

    For baseline jpeg files proposed conversion is straightforward as it's known when spectral stride for each component is ready.
    For progressive it's s bit tricky as progressive can define separate scans for each component so if(spectralEnd == 63) /* Convert spectral to color here */ won't work. There are two solutions:

    public Image<TPixel> Decode<TPixel>(BufferedReadStream stream, CancellationToken cancellationToken)
        where TPixel : unmanaged, IPixel<TPixel>
    {
        // this is still WIP but final variant would look somewhat like this
        var specificConverter = new SpectralToImageConverter<TPixel>(this.Configuration);
        this.spectralConverter = specificConverter;
    
        this.ParseStream(stream, cancellationToken: cancellationToken);
        this.InitExifProfile();
        this.InitIccProfile();
        this.InitIptcProfile();
        this.InitDerivedMetadataProperties();
    
        // this looks out of place to be honest
        if (/* This jpeg is progressive */)
        {
            specificConverter.ConvertFullScan();
        }
        
        return new Image<TPixel>(this.Configuration, this.Metadata, new[] { specificConverter.ImageFrame });
    }

    Another solution is to comit spectral data to the converter before returning from ParseStream() method.

    While both solutions look a bit 'ugly' it's the most performant way of checking are all scans done?. What do you think?

    P.S.
    Yes, Image<TPixel> creation looks ugly, I will open a separate PR for new ctor from single frame soon.

  13. antonfirsov commented on Jun 29, 2021

    @antonfirsov
    MemberAuthor

    @br3aker I like the plan with the pseudo Decode<TPixel> method. I don't think there is a way to avoid this complexity (or "ugliness"), since it comes from the jpeg standard itself.

  14. br3aker commented on Jun 29, 2021

    @br3aker
    Contributor

    @antonfirsov I will hide if-check in property getter for visual clarity then. Thanks for the responce!

  15. br3aker commented on Jul 7, 2021

    @br3aker
    Contributor

    A little update on this:

    I redid a lot of code and screwed something up and I couldn't find out why in a couple of hours + got a new idea which should be a little more understandable.

    First of all, resulting image whould be constructed from Buffer2D<TPixel> buffer. This would introduce a new Image<TPixel> constructor but it would be less intrusive. @antonfirsov not sure what to do in with #1680, implement pixel buffer ctor there or within this issue PR?

    Second, I've decided to implement this as an enumerable collection of spectral strides:

    // There would actually be some wrapping class for strides 
    // so it would store all components' spectral strides in a single object
    foreach(Buffer2D<Block8x8> spectralStride in scanDecoder)
    {
        // spectral -> vector4
        Buffer2D<Vector4> colorBuffer = ConvertFromSpectalToVector4(spectralStride);
    
        // vector4 -> TPixel
        PixelBuffer[i] = ConvertFromVector4<TPixel>(colorBuffer);
    }

    For every decoding mode it won't change anything (except there's won't be extra virtual call for each stride conversion) but for baseline dct it would allow to deffer stream parsing stride by stride.

  16. antonfirsov commented on Jul 8, 2021

    @antonfirsov
    MemberAuthor

    @antonfirsov not sure what to do in with #1680, implement pixel buffer ctor there or within this issue PR?

    If you have a working PR, that would be a good trigger to push things into a decision, otherwise it's just endless bike shadding API discussions 😄

    Second, I've decided to implement this as an enumerable collection of spectral strides.
    foreach(Buffer2D<Block8x8> spectralStride in scanDecoder)

    I wonder how does this work with ProcessStartOfScan? Does SOS kick of the code that iterates through the enumerable stuff, reading the stream further internally? What if there are more SOS-s?

  17. JimBobSquarePants commented on Jul 8, 2021

    @JimBobSquarePants
    Member

    What if there are more SOS-s?

    We definitely need to cater for that since that's how progressive jpegs work.

    I would open a draft PR where we can discuss the actual implementation.

  18. br3aker commented on Jul 8, 2021

    @br3aker
    Contributor

    I wonder how does this work with ProcessStartOfScan? Does SOS kick of the code that iterates through the enumerable stuff, reading the stream further internally? What if there are more SOS-s?

    That's not that hard to determine actually. Multiple SOS markers can exist only in:

    1. progressive jpegs - can be checked via SOF2 marker existence, current internal API already has it
    2. non-interleaved jpegs - can be checked if given scan component count is not equal to 'global' component count defined in SOF segment

    In other words:

    if (!this.Frame.IsProgressive && this.Frame.ComponentCount == scanComponentCount) 
    {
        // this SOS must be the only one, any extra is an error and can be checked after spectral decoding
        this.scanDecoder.Baseline = true;
    
        // we can return true to signal that we are ready for spectral conversion
        return true;
    }
    
    // decodes current partial scan to the pre-allocated spectral buffer
    this.scanDecoder.DecodeScan();
    
    // for more consistent behaviour we can actually evaluate if multi-sos jpeg is done
    // via spectralEnd == 63 for progressive jpegs 
    // via processedScans == this.Frame.ComponentCount for non-inteleaved baseline jpegs
    return lastScanCondition;

    There's a problem if given jpeg has anything after SOS except EOI and we can actually check even that - we can call ParseStream() one more time.

  19. br3aker commented on Jul 8, 2021

    @br3aker
    Contributor

    Nevermind, current architecture & code is not in a good shape for my plan, I'll try to work on it later. Priority right now:

    1. Working PR fixing this issue with least code change possible
    2. PR for refactoring (a lot of decoupling needed for decoder core & scan decoder) and micro-optimizations (there are some good places to cut off couple of ms)
    3. (possible) PR for Enumerable idea

    Sorry for this rapid change of ideas & messages, yesterday's discard of an almost working code knocked me hard. Will try to work out an implementation in a couple of days.

  20. JimBobSquarePants commented on Jul 9, 2021

    @JimBobSquarePants
    Member

    yesterday's discard of an almost working code knocked me hard

    No need to apologize and I feel you mate. At some point I need to replay months of optimization code I wrote for a zlib stream implementation because at some point I broke it but have no idea when. 😞

    Really looking forward to seeing what you come up with!

  21. br3aker commented on Jul 11, 2021

    @br3aker
    Contributor

    @antonfirsov @JimBobSquarePants sorry for bothering but I have a little problem

    PR is actually ready with almost all tests passing. The only problem is these tests for baseline jpegs:

    image

    These tests check if spectral data is equal to libjpeg spectral data of the given image. This approach is impossible as spectral data is discarded deep inside scan decoding process. Progressive and multi-scan baselines can be tested simply because they use the same technique as before the PR.

    Question: Do we even need to test spectral data? We can compare final colors which would be invalid if spectral is invalid. Yes, it's a couple layers 'higher' but right now it's impossible to test. Only if Enumerable approach will be implemented - spectral strides would be given out by virtual method so it would be possible to inject testing code.

    P.S.
    jpeg400jfif.jpg must fail - it's a known bug, don't mind it.

  22. antonfirsov commented on Jul 11, 2021

    @antonfirsov
    MemberAuthor

    We can compare final colors which would be invalid if spectral is invalid.

    The good thing about verifying spectral data is that the spectral intermediate result is exact, while for color conversion small deviations are allowed. It's important to be able to catch cases when a small difference is a result of a Huffman-decoding bug and not a floating point inaccuracy. We faced and fixed such issues while finalizing #274, and while I don't expect a refactor of that volume anytime soon, the tests can be still handy for Huffman decoding optimizations, so I prefer to keep them in long term.

    On the other hand, this should not block progress, so my recommendation is to temporarily disable them, and re-enable when the Enumerable refactor is done. @JimBobSquarePants agreed?

  23. br3aker commented on Jul 11, 2021

    @br3aker
    Contributor

    @antonfirsov skipping tests now seems like an easy plan but you know...

    I will slightly alter current PR architecture to enable testing without major changes so it won't rely on 'somewhat possible enumerable implementation in the future'.

  24. JimBobSquarePants commented on Jul 11, 2021

    @JimBobSquarePants
    Member

    @br3aker

    I will slightly alter current PR architecture to enable testing without major changes so it won't rely on 'somewhat possible enumerable implementation in the future'

    Not sure I follow what you mean here? Do you mean that this is not an issue now? I'm happy to temporarily disable baseline tests for now simply to see a difference.

  25. br3aker commented on Jul 11, 2021

    @br3aker
    Contributor

    Not sure I follow what you mean here? Do you mean that this is not an issue now? I'm happy to temporarily disable baseline tests for now simply to see a difference.

    It is a problem because I wanted to change as little code as possible for this PR so you guys won't spend too much time reviewing. These tests fix would result in a bigger change than necessary for PR to work. I will disable them and push a draft PR then.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions