美国心理学家 L. L. Thurstone 在他 1927 年的论文, “A 比较判断定律,” 中提出,当人们在多个选项中选择一个选项, 时,他们会选择对他们来说价值最高的选项,,即使他们无法为该选项指定一个特定的数字。

In his 1927 paper, “A law of comparative judgment,” the American psychologist L. L. Thurstone proposed that when people select one option among multiple alternatives, they are picking the one that has the highest value to them, even though they cannot assign a particular number to that choice. 

瑟斯顿是 “ 心理测量学” — 的先驱,该领域的前提是我们无法看到, 的心理过程, 可以被测量和量化。他 1927 年的论文为现在所谓的随机效用模型, 奠定了基础,该模型提供了一个数学框架来描述人类偏好— 信息,这些信息可以依赖, 依次, 对各种假设情况进行预测。

Thurstone was a pioneer of “psychometrics” — a field built upon the premise that mental processes, which we cannot see, can nevertheless be measured and quantified. His 1927 paper laid the groundwork for what are now called random utility models, which provide a mathematical framework for describing human preferences — information that can be relied upon, in turn, to make predictions about various hypothetical situations.

RUM, 当然, 在政府和工业界经常使用,其后果比选择热( 或冰) 饮料的后果要大得多。这些模型通常有助于预测人们在所谓的反事实 (“what-if”) 场景中会选择做什么,例如: 如果主干道因施工而关闭,他们将如何上班或上学? 他们将采取什么路线和交通方式? 或者, 如果一个城市突然收到 $.2 亿的意外之财, 这些资金应该如何使用进行支出以实现共同利益最大化?

RUMs, to be sure, are frequently used within government and industry in situations of far greater consequence than the selection of a hot (or iced) beverage. The models routinely facilitate predictions regarding what people will elect to do in so-called counterfactual (“what-if”) scenarios such as: How will they get to work or school if a major thoroughfare is shut down for construction? What routes and modes of transport will they take? Or, if a city suddenly receives a windfall of $20 million, how should those funds be disbursed to maximize the common good?

鉴于 RUM 已经陪伴我们近 100 年了, 随着时间的推移而变得更加复杂, 人们可能会想象, 在这个阶段, 几乎没有改进的空间。 , 然而, 并非如此。

Given that RUMs have been with us for almost 100 years, growing in sophistication over time, one might imagine that, at this stage, there would be little room for improvement. That, however, is not the case. 

的 小组的发现源于 , 部分 , 中 RUM 的估计方式存在缺陷,,这种缺陷自 Thurstone 时代以来一直存在。估计模型所依据的数据主要来自所谓的成对比较: 在项目 A 和 B 之间进行选择 — 是否与 Netflix 上的电影有关, Amazon.com 上的竞争产品, Google 上发布的新闻报道, 等等 — 您会选择哪一个? 这种方法如此普遍的原因之一, 解释了 Daskalakis, “ 为您从单个项目中获得的收益分配一个精确的数字分数 ,(例如 4.37,)是非常困难的。而比较两件事,并决定你更喜欢哪一个,在认知上更容易做到。” 但他补充说,这其中存在着问题,。 “用这种评估人们的方式’的偏好,一次只看两件事,不可能找到众多选择之间的相关性。”

The group的 findings stem, in part, from a deficiency in the way RUMs are commonly estimated in practice, which has persisted since the days of Thurstone. The data upon which the models are estimated have been largely drawn from so-called pairwise-comparisons: In a choice between items A and B — whether it pertains to movies on Netflix, competing products on Amazon.com, news stories posted on Google, and so forth — which one would you pick? One reason this approach has been so pervasive, explains Daskalakis, is that “assigning a precise numerical score, such as 4.37, to the benefit you get from a single item is very hard. Whereas comparing two things, and deciding which one you like better, is cognitively much easier to do.” But therein lies the rub, he adds. “With this way of assessing people的 preferences, looking at just two things at a time, it is impossible to find correlations between the numerous choices.”

应用 RUM 的标准方法假设从 A 和 B 派生的实用程序是独立的,,但它们可能, 事实上, 是链接的,,了解这一点很重要。如果竞选选举公职的人发现潜在选民支持枪支管制,,例如,,那么同一个人也有可能支持政府资助的儿童保育。同样,独立电影迷也可能偏爱外国电影,但对好莱坞动作大片不太热衷。 “如果数字平台对此类相关性的存在视而不见,,它将无法非常准确地估计偏好,” Daskalakis 指出。 “如果 Netflix 定期向您播放各种您不’不关心的电影,,您可能会退出并取消订阅。”

The standard way of applying RUMs assumes that the utilities derived from A and B are independent, but they may, in fact, be linked, and that would be important to know. If someone campaigning for elective office finds out that a potential voter favors gun control, for instance, there is a reasonable chance that same person also favors government-sponsored child care. Similarly, a fan of independent movies might also be partial to foreign films, but less enthusiastic about Hollywood action blockbusters. “If a digital platform has a blind eye to the existence of such correlations, it will not be able to estimate preferences very accurately,” Daskalakis notes. “And if Netflix regularly shows you an assortment of movies you don’t care about, you might sign off and cancel your subscription.”

麻省理工学院的团队证明,仅通过双向比较不可能获得相关性信息。然而,当大量人按照自己的偏好顺序对三种替代方案进行评分时,可以看出, 之间的相关性。相同的信息也可以从三局两胜和两局两胜选择的组合中获得。实际上, Mohammadpour 解释说, “ 你会让一群人对三个项目进行排名。然后,您可以利用我们开发的方法将这些单独的结果合并到一个大模型中,该模型可以为我们提供全局信息。”

The MIT team proved that it is impossible to get information about correlations from two-way comparisons alone. Correlations can be discerned, however, when large numbers of people rate three alternatives in their order of preference. The same information can also be obtained from a combination of best-of-three and best-of-two choices. In practice, Mohammadpour explains, “you would get a bunch of people to rank three items. You could then utilize the method we developed for merging those individual results into one big model that can provide us with the big picture.”

根据 Farina, 的说法,他们的研究工作, 集中在 RUM, 的计算方面,设计可以提取偏好信息的算法,并计算出需要多少数据才能完成此操作,或者, 相当于, 需要运行多少次实验。他说, 的好消息, 是有效的算法, 确实, 可以用于此目的。所需的实验数量不会随着目录或数据库中 正在审查的项目数量呈指数增长。

Their research effort, according to Farina, is focused on the computational side of RUMs, devising algorithms that can extract preference information and figuring out how much data is needed to do so or, equivalently, how many experiments need to be run. The good news, he says, is that efficient algorithms are, indeed, possible for this purpose. The requisite number of experiments does not grow exponentially with the number of items in the catalog or database that的 under review.

“这篇论文提供了一个关键的突破,”评论Emma Frejinger,蒙特利尔大学的计算机科学家。 “它从数学上证明了传统数据收集失败的原因,并证明只需询问用户三个中最好的[选择]即可解锁准确训练这些强大模型的能力。这一发现为收集更好的数据以推动更准确的优化提供了非常实用的路线图。”

“This paper provides a crucial breakthrough,” comments Emma Frejinger, a computer scientist at the University of Montreal. “It mathematically proves why traditional data collection fails and demonstrates that simply asking users for their best-of-three [choices] unlocks the ability to accurately train these powerful models. This finding provides a highly practical roadmap for collecting better data to drive more accurate optimizations.”

“ 建筑实用模型仍将是一个非常活跃的领域,” Daskalakis 坚持认为。 “正如自 1990 年代末以来 RUM 对互联网经济至关重要一样,,它们现在是,,并将继续成为, 对于未来人工智能模型的协调至关重要。” 更重要的是, 他补充道, “RUM 在大型语言模型的商业可行性和实用性方面发挥着核心作用[LLMs].” 在训练期间,,人们通常被要求对这些 LLM, 的各种候选输出进行排名,从中模型可以更好地了解文本 — 的类型,即语气, 风格 , 和内容 — 是首选。

“Building utility models is going to remain a very active area,” Daskalakis insists. “Just as RUMs have been critical to the internet economy since the late 1990s, they are, and will remain to be, critical to the alignment of AI models going forward.” More importantly, he adds, “RUMs play a central role in the commercial viability and usefulness of large language models [LLMs].” During the training period, people are typically asked to rank the various candidate outputs of these LLMs, from which the models can gain a better sense as to the kind of text — in terms of tone, style, and content — that is preferred. 

鉴于我们’不断地“在如此多的不同领域中被大量的选择所包围,” Daskalakis 说, “你不可能要求人们传达他们对所有可能情况的所有个人偏好。因此,您可以做的是建立一个模型来预测人们对不同可能结果的看法。并且您必须在迭代过程中不断改进和更新模型,直到, 希望, 您可以做出良好的预测。”

Given that we’re constantly “besieged with a vast sea of options in so many different domains,” Daskalakis says, “you cannot possibly ask people to communicate all their personal preferences for all possible scenarios. So what you can do instead is build a model that predicts what people think about the different possible outcomes. And you have to keep improving and updating your model in an iterative process until, hopefully, you can make good predictions.”