Table of Contents
Fetching ...

FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation

Kristýna Onderková, Ondřej Plátek, Zdeněk Kasner, Ondřej Dušek

TL;DR

FreshTab tackles evaluation contamination in table-to-text generation by auto-generating up-to-date benchmarks from Wikipedia tables that exceed LLM knowledge cutoffs. It introduces domain labels and multilingual data, uses reference-free and human evaluations, and compares against legacy benchmarks to show FreshTab's greater challenge and domain sensitivity. The study reveals misalignment between automatic metrics and human judgments, emphasizing the value of LLM-based judging, and demonstrates practical feasibility with monthly updates and non-English data. Overall, FreshTab offers a scalable, domain-aware, multilingual framework for more reliable table-to-text evaluation in evolving data landscapes.

Abstract

Table-to-text generation (insight generation from tables) is a challenging task that requires precision in analyzing the data. In addition, the evaluation of existing benchmarks is affected by contamination of Large Language Model (LLM) training data as well as domain imbalance. We introduce FreshTab, an on-the-fly table-to-text benchmark generation from Wikipedia, to combat the LLM data contamination problem and enable domain-sensitive evaluation. While non-English table-to-text datasets are limited, FreshTab collects datasets in different languages on demand (we experiment with German, Russian and French in addition to English). We find that insights generated by LLMs from recent tables collected by our method appear clearly worse by automatic metrics, but this does not translate into LLM and human evaluations. Domain effects are visible in all evaluations, showing that a~domain-balanced benchmark is more challenging.

FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation

TL;DR

FreshTab tackles evaluation contamination in table-to-text generation by auto-generating up-to-date benchmarks from Wikipedia tables that exceed LLM knowledge cutoffs. It introduces domain labels and multilingual data, uses reference-free and human evaluations, and compares against legacy benchmarks to show FreshTab's greater challenge and domain sensitivity. The study reveals misalignment between automatic metrics and human judgments, emphasizing the value of LLM-based judging, and demonstrates practical feasibility with monthly updates and non-English data. Overall, FreshTab offers a scalable, domain-aware, multilingual framework for more reliable table-to-text evaluation in evolving data landscapes.

Abstract

Table-to-text generation (insight generation from tables) is a challenging task that requires precision in analyzing the data. In addition, the evaluation of existing benchmarks is affected by contamination of Large Language Model (LLM) training data as well as domain imbalance. We introduce FreshTab, an on-the-fly table-to-text benchmark generation from Wikipedia, to combat the LLM data contamination problem and enable domain-sensitive evaluation. While non-English table-to-text datasets are limited, FreshTab collects datasets in different languages on demand (we experiment with German, Russian and French in addition to English). We find that insights generated by LLMs from recent tables collected by our method appear clearly worse by automatic metrics, but this does not translate into LLM and human evaluations. Domain effects are visible in all evaluations, showing that a~domain-balanced benchmark is more challenging.
Paper Structure (26 sections, 9 figures, 9 tables)

This paper contains 26 sections, 9 figures, 9 tables.

Figures (9)

  • Figure 1: Schema of the FreshTab method
  • Figure 2: TAPEX on LoTNLG vs. FreshTab.2-5/25.en.lotvs. FreshTab.2-5/25.en.diverse
  • Figure 3: Llama-as-a-judge on LoTNLG vs. FreshTab.2-5/25.en.lot vs. FreshTab.2-5/25.en.diverse
  • Figure 4: Total number of errors found in human evaluation by model and benchmark
  • Figure 5: TAPEX and Llama as a judge on FreshTab.2-5/25.en.diverse by domain.
  • ...and 4 more figures