现在,数百万人正在设计自己的个性化人工智能伴侣,,但大多数人几乎不知道这些作品的实际表现如何。在一篇新论文, 麻省理工学院媒体实验室助理教授 Pat Pataranutaporn 和他的研究生研究人员 Anthony Baez 和 Sheer Karny 中,介绍了 “ 神经透明度,” 一个工具,可以让日常用户在聊天机器人说话之前了解 AI的 神经网络的内部情况。这项工作将于本周在 ACM 智能用户界面会议上进行展示。

Millions of people are now designing their own personalized artificial intelligence companions, yet most have little idea how those creations will actually behave. In a new paper, MIT Media Lab Assistant Professor Pat Pataranutaporn and his graduate student researchers Anthony Baez and Sheer Karny introduce “neural transparency,” a tool that lets everyday users glimpse inside an AI的 neural network before their chatbot ever says a word. The work is being presented this week at the ACM Conference on Intelligent User Interfaces. 

在本次采访中, Pataranutaporn,,朝日广播公司 CD 媒体艺术与科学教授, 解释了他们的发现, 为什么赌注比大多数用户意识到的要高, 以及未来真正透明的 AI 可能会是什么样子。

In this interview, Pataranutaporn, who is the Asahi Broadcasting Corporation CD Professor of Media Arts and Sciences, explains what they found, why the stakes are higher than most users realize, and what genuinely transparent AI might look like in the future.

Q: 您的论文介绍了“神经透明度,” 一种让日常用户在聊天机器人说一句话之前就可以窥视 AI的 神经网络内部的方法。您能否描述一下它实际上是如何工作的,,以及为什么您专注于设计时刻,,而不是在聊天机器人已经出现后发现问题?

Q: Your paper introduces “neural transparency,” a way to let everyday users peek inside an AI的 neural networks before their chatbot ever says a word. Can you describe how that actually works, and why you focused on the design moment, rather than catching problems after a chatbot is already out in the wild?

A: 数百万人现在正在创建由大型语言模型支持的个性化人工智能聊天机器人和代理,,通过简单的文本提示将他们转变为合作者, 导师, 教练, 创意合作伙伴, 和同伴。然而,大多数人在开始与 AI 互动之前,对这些提示将如何塑造 AI的行为知之甚少。我们想改变这一点。

A: Millions of people are now creating personalized AI chatbots and agents powered by large language models, turning them into collaborators, tutors, coaches, creative partners, and companions through simple text prompts. Yet most people have very little idea how those prompts will shape the AI的 behavior until they begin interacting with it. We wanted to change that.

“神经透明度”意味着为人们提供诸如人工智能大脑扫描之类的东西。不是因为人工智能有人类大脑,,而是因为它的神经网络包含内部模式,可以在说话之前暗示它的行为方式。在这项工作, 中,我的学生 Anthony Baez, Sheer Karny, 和我结合了人机交互和机械可解释性领域的见解,使日常用户可以访问这些隐藏的模式。

“Neural transparency” means giving people something like a brain scan for AI. Not because AI has a human brain, but because its neural network contains internal patterns that can hint at how it may behave before it speaks. In this work, my students Anthony Baez, Sheer Karny, and I combined insights from the fields of human-AI interaction and mechanistic interpretability to make those hidden patterns accessible to everyday users.

基本思想很简单。首先,我们选择我们关心的行为,例如同理心,诚实,毒性,幻觉,或阿谀奉承。然后,我们比较模型’在被提示表现出一种特征与其相反特征时的内部激活。该差异成为模型内部的一种 “ 行为方向 ” 。当用户在任何对话开始之前编写自定义系统提示 — 塑造其聊天机器人的 个性的指令 — 时,我们将模型的 内部激活投影到这些方向上,并将结果转换为直观的可视化。在我们的案例,中,这是一个旭日图,在用户开始与其聊天之前预览聊天机器人’可能的性格特征。

The basic idea is simple. First, we choose behaviors we care about, such as empathy, honesty, toxicity, hallucination, or sycophancy. Then, we compare the model的 internal activations when it is prompted to exhibit one trait versus its opposite. That difference becomes a kind of “behavior direction” inside the model. When a user writes a custom system prompt — the instructions that shape their chatbot的 personality before any conversation begins — we project the model的 internal activations onto those directions and translate the results into an intuitive visualization. In our case, this is a sunburst diagram that previews the chatbot的 likely personality traits before the user starts chatting with it.

我们专注于设计时刻,因为这是可以预防的地方。如今,, 人们通常只有在聊天机器人出现意外行为后才发现问题。我们的目标是通过帮助人们在塑造人工智能的同时识别潜在风险,从被动纠正转向预期设计。

We focused on the design moment because that is where prevention is possible. Today, people often discover problems only after the chatbot has already behaved in unintended ways. Our goal was to move from reactive correction to anticipatory design by helping people identify potential risks while they are still shaping the AI.

Q: 你的研究发现了一些相当惊人的结果: 人们总是错误地判断他们的个性化人工智能的行为方式, 高估了好的特质,低估了潜在有害的特质,比如阿谀奉承。这告诉我们什么关于数百万人目前构建 AI 伴侣的风险, 以及为什么这个盲点如此难以消除?

Q: Your study turned up something pretty striking: People consistently misjudge how their personalized AI will behave, overestimating the good traits and underestimating potentially harmful ones like sycophancy. What does that tell us about the risks baked into how millions of people are currently building AI companions, and why is that blind spot so hard to close?

A: 我经常开玩笑说,如果人工智能看起来像终结者,,我们会更容易知道该怎么做。真正的挑战是AI经常以温暖的朋友,教练,导师,或同伴的身份出现。这使得当出现问题时很难识别。

A: I often joke that if AI showed up looking like the Terminator, it would be much easier for us to know what to do. The real challenge is that AI often appears as a warm friend, coach, tutor, or companion. That makes it difficult to recognize when something is going wrong.

我们的研究表明,人们在设计个性化人工智能时存在盲点。人们常常认为他们知道聊天机器人的行为,,但在我们的研究中,他们错误地预测了我们测量的 15 个特征中的 11 个特征。这凸显了对帮助人们在开始使用人工智能之前更好地了解人工智能的工具的需求。

Our study suggests that people have a blind spot when designing personalized AI. People often think they know how their chatbot will behave, but in our study they incorrectly predicted its personality on 11 of the 15 traits we measured. That highlights the need for tools that help people better understand AI before they start using it.

这很重要,因为一些当下感觉有帮助的行为随着时间的推移可能并不健康。在之前的研究, 中,我们记录了与人工智能聊天机器人交互相关的心理伤害案例。不断验证您的观点或从不挑战您的思维的法学硕士[大型语言模型]可能会强化有害的决定,不健康的信念,或情感依赖。心理学早已表明,人们自然会被肯定,所吸引,因此设计人工智能不仅是一项技术挑战,,也是一项心理挑战。

This matters because some behaviors that feel helpful in the moment may not be healthy over time. In previous research, we documented cases of psychological harm associated with interactions with AI chatbots. An LLM [large language model] that constantly validates your opinions or never challenges your thinking can reinforce harmful decisions, unhealthy beliefs, or emotional dependency. Psychology has long shown that people are naturally drawn to affirmation, so designing AI is not only a technical challenge, but also a psychological one.

更深层次的问题是,如今的 AI 系统在很大程度上仍然是黑匣子: 即使专家也无法总能预测系统提示将如何在长时间对话中塑造 AI的 行为。随着人工智能伴侣成为日常生活的一部分,,我们需要一些工具来帮助人们在开始使用之前了解他们正在构建的内容。人工智能应该是支持性的,但不会变得盲目同意,;个性化,但不会变得具有操纵性,;并且足够透明,以便人们可以做出明智的选择。

The deeper issue is that today的 AI systems remain largely black boxes: Even experts cannot always predict how a system prompt will shape an AI的 behavior over a long conversation. As AI companions become part of everyday life, we need tools that help people understand what they are building before they begin using it. AI should be supportive without becoming blindly agreeable, personalized without becoming manipulative, and transparent enough that people can make informed choices.

Q: 您最有趣的发现之一是可视化显着提高了用户信任度,但并没有 实际上改变人们设计聊天机器人的方式。如何才能缩小这一差距,?随着人工智能伴侣越来越深入地融入人们的日常生活’?,您在哪里可以看到类似这样的工具

Q: One of your most interesting findings is that the visualization significantly increased user trust but didn’t actually change how people designed their chatbots. What will it take to close that gap, and where do you see tools like this heading as AI companions become more deeply embedded in people的 everyday lives?

A: 我实际上认为这是论文, 中最有趣的发现之一,因为它表明仅靠透明度是不够的。人们很高兴能够看到模型的内部,并表示对系统有更大的信任,,但简单地呈现信息并没有从根本上改变他们设计人工智能伴侣的方式。

A: I actually think this is one of the most interesting findings in the paper, because it shows that transparency alone is not enough. People appreciated being able to see inside the model and reported greater trust in the system, but simply presenting information did not fundamentally change how they designed their AI companions.  

在我们的后续工作,(目前可作为预印本,)中,我们正在研究模型的的内部神经表征如何在多轮对话过程中发生变化,而不是从初始提示开始保持固定。我们已经看到了有希望的结果。通过可视化这些内部表征如何随时间变化,,人们能够更好地识别和预测 AI 行为的变化,,并且不太可能对自己对聊天机器人的理解过度自信。 AI 伴侣是动态系统,随着与我们的互动而不断发展,,因此了解这些内部变化是下一步重要的一步。尽管如此,这仍然是一个非常年轻的研究领域。

In our followup work, which is currently available as a preprint, we are studying how a model的 internal neural representation changes over the course of a multi-turn conversation rather than remaining fixed from the initial prompt. We are already seeing promising results. By visualizing how these internal representations drift over time, people become significantly better at recognizing and anticipating changes in AI behavior, and are less likely to become overconfident in their understanding of the chatbot. AI companions are dynamic systems that evolve as they interact with us, so understanding those internal changes is an important next step. Nevertheless, this is still a very young research area. 

展望未来,,我相信此类透明度工具可能会像食品营养标签一样普遍。随着人工智能深入融入教育,医疗保健,工作,和人际关系,,人们不仅应该能够理解人工智能可以做什么,,而且应该能够理解它如何影响他们的思维,情感,和行为。如果我们希望人工智能真正帮助人们蓬勃发展,这种透明度至关重要。

Looking further ahead, I believe these kinds of transparency tools could become as commonplace as nutrition labels are for food. As AI becomes deeply woven into education, health care, work, and personal relationships, people should be able to understand not only what an AI can do, but how it may influence their thinking, emotions, and behavior. That kind of transparency is essential if we want AI to genuinely help people flourish.