Table of Contents
Fetching ...

How to Sell High-Dimensional Data Optimally

Andrew Li, R. Ravi, Karan Singh, Zihong Yi, Weizhong Zhang

TL;DR

The paper tackles selling high-dimensional proprietary data by modeling information as a menu of statistical experiments that improve buyers' decision quality. It develops a sampling-based algorithm that yields near-optimal menus with runtime and sample complexity independent of the (potentially exponential) state space, and analyzes a high-dimensional Gaussian special case to obtain concrete, efficiently computable menu designs. It shows that scalar Gaussian experiments suffice, provides an SDP to compute the revenue-maximizing menu, and identifies a separation condition on buyer preferences that enables full surplus extraction; it also proves that deterministic signaling can be optimal in high dimensions. These results offer scalable, principled tools for pricing data products in real-world markets and shed light on the value of offering differentiated data products versus attempting to reveal all information.

Abstract

Motivated by the problem of selling large, proprietary data, we consider an information pricing problem proposed by Bergemann et al. that involves a decision-making buyer and a monopolistic seller. The seller has access to the underlying state of the world that determines the utility of the various actions the buyer may take. Since the buyer gains greater utility through better decisions resulting from more accurate assessments of the state, the seller can therefore promise the buyer supplemental information at a price. To contend with the fact that the seller may not be perfectly informed about the buyer's private preferences (or utility), we frame the problem of designing a data product as one where the seller designs a revenue-maximizing menu of statistical experiments. Prior work by Cai et al. showed that an optimal menu can be found in time polynomial in the state space, whereas we observe that the state space is naturally exponential in the dimension of the data. We propose an algorithm which, given only sampling access to the state space, provably generates a near-optimal menu with a number of samples independent of the state space. We then analyze a special case of high-dimensional Gaussian data, showing that (a) it suffices to consider scalar Gaussian experiments, (b) the optimal menu of such experiments can be found efficiently via a semidefinite program, and (c) full surplus extraction occurs if and only if a natural separation condition holds on the set of potential preferences of the buyer.

How to Sell High-Dimensional Data Optimally

TL;DR

The paper tackles selling high-dimensional proprietary data by modeling information as a menu of statistical experiments that improve buyers' decision quality. It develops a sampling-based algorithm that yields near-optimal menus with runtime and sample complexity independent of the (potentially exponential) state space, and analyzes a high-dimensional Gaussian special case to obtain concrete, efficiently computable menu designs. It shows that scalar Gaussian experiments suffice, provides an SDP to compute the revenue-maximizing menu, and identifies a separation condition on buyer preferences that enables full surplus extraction; it also proves that deterministic signaling can be optimal in high dimensions. These results offer scalable, principled tools for pricing data products in real-world markets and shed light on the value of offering differentiated data products versus attempting to reveal all information.

Abstract

Motivated by the problem of selling large, proprietary data, we consider an information pricing problem proposed by Bergemann et al. that involves a decision-making buyer and a monopolistic seller. The seller has access to the underlying state of the world that determines the utility of the various actions the buyer may take. Since the buyer gains greater utility through better decisions resulting from more accurate assessments of the state, the seller can therefore promise the buyer supplemental information at a price. To contend with the fact that the seller may not be perfectly informed about the buyer's private preferences (or utility), we frame the problem of designing a data product as one where the seller designs a revenue-maximizing menu of statistical experiments. Prior work by Cai et al. showed that an optimal menu can be found in time polynomial in the state space, whereas we observe that the state space is naturally exponential in the dimension of the data. We propose an algorithm which, given only sampling access to the state space, provably generates a near-optimal menu with a number of samples independent of the state space. We then analyze a special case of high-dimensional Gaussian data, showing that (a) it suffices to consider scalar Gaussian experiments, (b) the optimal menu of such experiments can be found efficiently via a semidefinite program, and (c) full surplus extraction occurs if and only if a natural separation condition holds on the set of potential preferences of the buyer.
Paper Structure (19 sections, 12 theorems, 18 equations, 1 table, 1 algorithm)

This paper contains 19 sections, 12 theorems, 18 equations, 1 table, 1 algorithm.

Key Result

Proposition 1

For any setting of type-dependent preferences, $R_\textrm{one}/R_\textrm{menu}$ and and $R_\textrm{menu}/R_\textrm{full-info}$ lie in $[1/n,1]$. Furthermore, for any $\varepsilon>0$, there exists a setting in which $R_\textrm{one}/R_\textrm{menu}\leq 1/n+\varepsilon$ and $R_\textrm{menu} = R_\textr

Theorems & Definitions (21)

  • Proposition 1
  • Remark
  • Theorem 1
  • proof : Proof of Theorem \ref{['thm:main']}:
  • Lemma 1
  • Lemma 2
  • Lemma 3
  • proof
  • Proposition 2
  • proof
  • ...and 11 more