Skip to content

Support different de/serialization strategies #478

Description

@LaurentRDC

Currently, all messages are de/serialized via the Binary instance.

This is an implementation detail; there is strict required beyond the existence of two functions, encode :: a -> ByteString and decode :: ByteString -> Either String a. However, discussions in the DataHaskell group have raised the point that serialization for data science structures (e.g. dataframes) could be made more efficient by leveraging Parquet over the wire.

I'm thinking of an extension of the current Process monad, to support also a de/serialization strategy. This is strongly inspired by Servant's MimeRender and MimeUnrender classes. Consider:

newtype Process strategy a = Process ... -- Monad is unchanged, type has a new `strategy`

class Serialize strategy a where
    encode :: Proxy strategy -> a -> ByteString
    decode :: Proxy strategy -> ByteString -> Either String a

An example for JSON would be:

import Data.Aeson

data JSON

class (ToJSON a, FromJSON a) => Serialize JSON a where
    encode _ = encode
    decode _ = eitherDecodeStrict 

Then, for the actual message sending e.g. via send, the type would become:

send :: Serialize strategy a => ProcessId -> a -> Process strategy ()

Then, programs that are agnostic of serialization are of type Process strategy a (instead of Process a), while programs which require a specific serialization method, such as JSON, will be of type Process JSON a.

Activity

  1. added this to the Horizon milestone on Oct 12, 2025
  2. changed the title [-]Support different de/serialization methods[/-] [+]Support different de/serialization strategies[/+] on Oct 12, 2025
  3. robinp commented on Oct 17, 2025

    @robinp

    Not that I have any stake in this, but why not use a newtype wrapper on those types, and make the Binary instance for the wrapped type use the desired serialization? For example

    newtype UsingParquet a = UsingParquet a
    
    instance Parquetable a => Binary (UsingParquet a) where
         encode a = someParquetEncode a ...
    

    This sounds less intrusive and more explicit, for the slight but standard inconvenience of needing to wrap/unwrap.

  4. LaurentRDC commented on Oct 17, 2025

    @LaurentRDC
    MemberAuthor

    Yep that also works in principle!

    In your example, the function someParquetEncode is doing a lot of work. Some things are easy to replicate using Binary. I imagine that some things are not.

  5. mchav commented on Nov 13, 2025

    @mchav

    https://arrow.apache.org/docs/format/IPC.html

    Would be a good format to support. Parquet is mostly meant for data at rest. Parquet has a little more overhead since it serializes metadata and column statistics for table optimizations like predicate push down etc.

    I can help look into that as I work on arrow more generally.

  6. LaurentRDC commented on Nov 14, 2025

    @LaurentRDC
    MemberAuthor

    Thanks @mchav , I'll look into the IPC format.

    As @robinp points out, if we can reuse the Binary class for any format (including IPC, then there's no breaking change, which is nice.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions