Bridging Items and Language: A Transition Paradigm for Large Language Model-Based Recommendation

Xinyu Lin; Wenjie Wang; Yongqi Li; Fuli Feng; See-Kiong Ng; Tat-Seng Chua

Bridging Items and Language: A Transition Paradigm for Large Language Model-Based Recommendation

Xinyu Lin, Wenjie Wang, Yongqi Li, Fuli Feng, See-Kiong Ng, Tat-Seng Chua

TL;DR

This work tackles bridging the item space and language space for LLM-based recommendations by introducing TransRec, a transition paradigm that uses multi-facet item identifiers (ID, title, and attributes) and position-free constrained generation empowered by an FM-index. An aggregated grounding module ties generated identifiers to in-corpus items, enabling robust ranking and improved generalization, including strong few-shot and cold-start performance with large LLM backbones. Extensive experiments on Beauty, Toys, and Yelp datasets demonstrate TransRec's superiority over traditional and prior LLM-based methods, while ablation and hyper-parameter analyses reveal the critical roles of each facet and the grounding mechanism. The approach advances practical LLM-based recommendation by ensuring semantic richness, discriminative item representation, and reliable in-corpus grounding, with potential for automatic facet construction and enhanced grounding strategies in future work.

Abstract

Harnessing Large Language Models (LLMs) for recommendation is rapidly emerging, which relies on two fundamental steps to bridge the recommendation item space and the language space: 1) item indexing utilizes identifiers to represent items in the language space, and 2) generation grounding associates LLMs' generated token sequences to in-corpus items. However, previous methods exhibit inherent limitations in the two steps. Existing ID-based identifiers (e.g., numeric IDs) and description-based identifiers (e.g., titles) either lose semantics or lack adequate distinctiveness. Moreover, prior generation grounding methods might generate invalid identifiers, thus misaligning with in-corpus items. To address these issues, we propose a novel Transition paradigm for LLM-based Recommender (named TransRec) to bridge items and language. Specifically, TransRec presents multi-facet identifiers, which simultaneously incorporate ID, title, and attribute for item indexing to pursue both distinctiveness and semantics. Additionally, we introduce a specialized data structure for TransRec to ensure generating valid identifiers only and utilize substring indexing to encourage LLMs to generate from any position of identifiers. Lastly, TransRec presents an aggregated grounding module to leverage generated multi-facet identifiers to rank in-corpus items efficiently. We instantiate TransRec on two backbone models, BART-large and LLaMA-7B. Extensive results on three real-world datasets under diverse settings validate the superiority of TransRec.

Bridging Items and Language: A Transition Paradigm for Large Language Model-Based Recommendation

TL;DR

Abstract

Paper Structure (34 sections, 7 equations, 9 figures, 9 tables)

This paper contains 34 sections, 7 equations, 9 figures, 9 tables.

Introduction
Preliminary
Instruction Tuning
Generation Grounding
Method
Multi-facet Item Indexing
Multi-facet Identifier
Data Reconstruction
Multi-facet Generation Grounding
Position-free Constrained Generation
Aggregated Grounding
Experiments
Experimental Settings
Datasets
Baselines
...and 19 more sections

Figures (9)

Figure 1: Illustration of the two pivotal steps for LLM-based recommenders: item indexing and generation grounding.
Figure 2: Overview of TransRec. Item indexing assigns each item a multi-facet identifier. For generation grounding, TransRec generates a set of identifiers in each facet and then grounds them to in-corpus items for ranking.
Figure 3: Illustration of instruction tuning of LLMs. (a) depicts the conversion from recommendation data to instruction data; (b) presents the optimization of LLMs based on the instruction data.
Figure 4: Illustration of the reconstructed data based on the multi-facet identifiers. The bold texts in black refer to the user's historical interactions.
Figure 5: Demonstration of the generation grounding step in TransRec. Red, blue, and green denote the facets of ID, title, and attribute, respectively.
...and 4 more figures

Bridging Items and Language: A Transition Paradigm for Large Language Model-Based Recommendation

TL;DR

Abstract

Bridging Items and Language: A Transition Paradigm for Large Language Model-Based Recommendation

Authors

TL;DR

Abstract

Table of Contents

Figures (9)