Learning with Structure: Computing Consistent Subsets on Structurally-Regular Graphs

Aritra Banik; Mano Prakash Parthasarathi; Venkatesh Raman; Diya Roy; Abhishek Sahu

Paper

Learning with Structure: Computing Consistent Subsets on Structurally-Regular Graphs

Abstract

The Minimum Consistent Subset (MCS) problem arises naturally in the context of supervised clustering and instance selection. In supervised clustering, one aims to infer a meaningful partitioning of data using a small labeled subset. However, the sheer volume of training data in modern applications poses a significant computational challenge. The MCS problem formalizes this goal: given a labeled dataset

in a metric space, the task is to compute a smallest subset

such that every point in

shares its label with at least one of its nearest neighbors in

. Recently, the MCS problem has been extended to graph metrics, where distances are defined by shortest paths. Prior work has shown that MCS remains NP-hard even on simple graph classes like trees, though an algorithm with runtime

is known for trees, where

is the number of colors and

the number of vertices. This raises the challenge of identifying graph classes that admit algorithms efficient in both

and

. In this work, we study the Minimum Consistent Subset problem on graphs, focusing on two well-established measures: the vertex cover number (

) and the neighborhood diversity (

). We develop an algorithm with running time

, and another algorithm with runtime

. In the language of parameterized complexity, this implies that MCS is fixed-parameter tractable (FPT) parameterized by the vertex cover number and the neighborhood diversity. Notably, our algorithms remain efficient for arbitrarily many colors, as their complexity is polynomially dependent on the number of colors.