parquet vs csv

GravitySpoiled@lemmy.ml · 6 months ago

parquet vs csv

The Hobbyist · 6 months ago

In the deep learning community, I know of someone using parquet for the dataset and annotations. It allows you to select which data you want to retrieve from the dataset and stream only those, and nothing else. It is a rather effective method for that if you have many different annotations for different use cases and want to be able to select only the ones you need for your application.

demesisx@infosec.pub · 6 months ago

How does this differ from graphQL?

ma343@beehaw.org · 6 months ago

Graphql is a protocol for interacting with a remote system, parquet is about having a local file that you can index and retrieve data from in a more efficient way. It’s especially useful when the data has a fairly well defined structure but may be large enough that you can’t or don’t want to bring it all into memory. They’re similar concepts, but different applications

demesisx@infosec.pub · 6 months ago

Thank you!

djnattyp@lemmy.world · 6 months ago

Parquet is a storage format; graphQL is a query language/transmission strategy.