Building a Macedonian Recipe Dataset: Collection, Parsing, and Comparative Analysis
Darko Sasanski, Dimitar Peshevski, Riste Stojanov, Dimitar Trajanov
TL;DR
The paper tackles the scarcity of Macedonian recipes in computational gastronomy by constructing the first structured Macedonian recipe dataset through web scraping and a robust parsing pipeline. It analyzes ingredient usage and co-occurrence with measures such as Pointwise Mutual Information (PMI) and Lift, comparing the Macedonian corpus to a subset of Recipe1M+ to reveal cultural coherence versus international diversity. The dataset comprises 36,237 recipes from three Macedonian sites and demonstrates strong domain-specific associations, particularly around baking and local dairy ingredients. This work provides a valuable resource for cultural preservation, nutrition analysis, and authentic recipe generation in a low-resource language, and it offers a replicable methodology for building language-focused culinary datasets.
Abstract
Computational gastronomy increasingly relies on diverse, high-quality recipe datasets to capture regional culinary traditions. Although there are large-scale collections for major languages, Macedonian recipes remain under-represented in digital research. In this work, we present the first systematic effort to construct a Macedonian recipe dataset through web scraping and structured parsing. We address challenges in processing heterogeneous ingredient descriptions, including unit, quantity, and descriptor normalization. An exploratory analysis of ingredient frequency and co-occurrence patterns, using measures such as Pointwise Mutual Information and Lift score, highlights distinctive ingredient combinations that characterize Macedonian cuisine. The resulting dataset contributes a new resource for studying food culture in underrepresented languages and offers insights into the unique patterns of Macedonian culinary tradition.
