Patufet is a collection of open Catalan datasets assembled for pretraining and fine-tuning language models. It grew alongside Cucafera: before training a local-language model, we had to obtain the data.

The collection includes educational text, stories, code, prompts, conversations, and instruction data. Publishing the datasets separately makes the pipeline reusable and lets others inspect, adapt, and question the material that shaped the model.

Patufet is available as a public resource on Hugging Face.