BTC $84,012.68 -0.89%
ETH $2,697.56 +0.03%
BNB $769.85 -1.11%
XRP $1.51 -1.67%
SOL $119.96 -2.59%
TRX $0.3363 +0.59%
DOGE $0.0949 -2.82%
ADA $0.2494 -3.00%
BCH $310.81 -7.49%
LINK $15.25 +7.86%
HYPE $88.75 -3.54%
AAVE $149.22 -4.21%
SUI $1.15 -8.14%
XLM $0.2305 +5.99%
ZEC $1,512.57 -6.17%
AAPL $339.28 -0.29%
AMZN $246.27 -1.52%
GOOGL $342.46 -0.42%
MSFT $511.34 -1.14%
META $720.84 -3.87%
NVDA $229.71 +1.95%
TSLA $358.34 -4.00%
SNDK $1,712.55 -3.96%
INTC $115.42 -8.11%
SPCX $147.33 -1.02%
MU $1,053.67 -3.95%
AMD $607.16 -4.16%
BTC $84,012.68 -0.89%
ETH $2,697.56 +0.03%
BNB $769.85 -1.11%
XRP $1.51 -1.67%
SOL $119.96 -2.59%
TRX $0.3363 +0.59%
DOGE $0.0949 -2.82%
ADA $0.2494 -3.00%
BCH $310.81 -7.49%
LINK $15.25 +7.86%
HYPE $88.75 -3.54%
AAVE $149.22 -4.21%
SUI $1.15 -8.14%
XLM $0.2305 +5.99%
ZEC $1,512.57 -6.17%
AAPL $339.28 -0.29%
AMZN $246.27 -1.52%
GOOGL $342.46 -0.42%
MSFT $511.34 -1.14%
META $720.84 -3.87%
NVDA $229.71 +1.95%
TSLA $358.34 -4.00%
SNDK $1,712.55 -3.96%
INTC $115.42 -8.11%
SPCX $147.33 -1.02%
MU $1,053.67 -3.95%
AMD $607.16 -4.16%

Zuckerberg Talks Muse: AI Agent Explosion, "Metaverse, Smart Glasses, and Large Models" Three Major Bets Converge

Core Viewpoint
Summary: Muse launched two weeks ago and has already reached millions of users, with Mark Zuckerberg stating, "It's a home run right out of the gate." This personal AI Agent is about to be fully integrated with Ray-Ban smart glasses, betting on the convergence of the metaverse, smart glasses, and large models. In addition, he rarely admitted that the failure of Llama 4 was "the most terrifying moment" and revealed that Meta is building a 5-gigawatt computing cluster.
Wall Street Journal
2026-09-27 23:36:54
Muse launched two weeks ago and has already reached millions of users, with Mark Zuckerberg stating, "It's a home run right out of the gate." This personal AI Agent is about to be fully integrated with Ray-Ban smart glasses, betting on the convergence of the metaverse, smart glasses, and large models. In addition, he rarely admitted that the failure of Llama 4 was "the most terrifying moment" and revealed that Meta is building a 5-gigawatt computing cluster.

Author: Long Yue, Wall Street Insights

Three years ago, Meta CEO Mark Zuckerberg mentioned on the Joe Rogan show that one day people would wear glasses, and AI Agents would emerge accordingly. Today, this scene may be unfolding.

Meta's personal AI Agent—Muse—has reached millions of users within two weeks of its launch. In an interview on September 25, Zuckerberg stated, "We only encounter this kind of situation every few years." He characterized the early response to Muse as "a home run right out of the gate"—a rarity in Meta's product history.

At the same time, he announced that Muse will be integrated into the entire line of Ray-Ban smart glasses, allowing users to call their personal AI Agent directly with a custom wake word instead of saying "Hey Meta."

For years, the metaverse, smart glasses, and large models, which have been viewed externally as three independent bets, are now accelerating convergence at the product level within Meta.

In this interview, Zuckerberg also admitted that the failure of Llama 4 was "the scariest moment" and revealed that Meta is building a 5-gigawatt training cluster, believing that sufficiently large computing power could "brute force" AGI.

Zuckerberg Talks Muse: AI Agent Explosion,

Why Muse Can Be "A Home Run Right Out of the Gate"

The starting point for Muse was when Zuckerberg and team members Nat and Alex sat together and "pieced together" an early version using open-source tools at home.

"We realized this was a magical experience," Zuckerberg said, "If we could make it accessible to anyone—without needing to buy a Mac Mini or tinker with terminals—then this would be something billions of people would want to use."

This judgment drove the entire subsequent research and development logic: not just training models, but fully self-developing the "heartbeat" mechanism of the Agent, from the model and scaffolding to the Agent itself—where the Agent would periodically wake up, check user goals, and proactively push tasks forward.

Data after the launch confirmed this judgment. In two weeks, millions of users.

Personal AI: Focus on Execution, Memory, and "Judgment"

Regarding the difference between personal AI and general AI, Zuckerberg summarized the current competitive focus as making models better Agents.

"The biggest thing over the past year has basically been programming Agents," he said. Programming capability is crucial because even if users do not write code directly, "your Muse will be writing code for you in the background to accomplish various tasks."

However, he believes that personal Agents cannot only have coding or task execution capabilities; they also need to understand privacy boundaries and social contexts.

Zuckerberg gave an example: when users ask Muse to book a restaurant, the Agent might know about the user's allergies or pregnancy, but that does not mean this information should be automatically disclosed. "You would want it to complete the task while disclosing as little information as possible."

He stated that this capability falls under the "basic social skills or common sense" that humans typically possess, but many companies have not included it in their model capabilities if they are only training coding Agents.

"We are training the entire model, not just taking someone else's ready-made model and building a framework and Agent around it," Zuckerberg said. Meta is also adding features like memory, task tracking, virtual character animations, and real-time voice interaction at the product level.

In terms of personalization, he stated that Meta does not believe AI should have only one fixed personality. "Many labs are focusing on how to tune personality to the right state, but I never believed there is only one correct answer for personality."

He mentioned that Muse allows users to modify their avatar, voice, and fundamental personality, with the model designed to have high controllability. "You can define how you want to interact with it, which is important for personal Agents."

Every Agent Has Its Own "Computer"

One of the core infrastructures that distinguishes Muse from other AI products is that Meta has equipped each Agent with a separate Secure Virtual Machine (Secure VM).

Zuckerberg explained why this is necessary:

Your Agent will know a lot of sensitive information about you. We do not believe this information should be mixed in a pool with everyone else's data.

He likened the Muse Secure VM to "the computer that only belongs to you under your desk"—user data is encrypted and stored, and sensitive credentials like passwords are managed through an independent security module, which the Agent itself cannot directly read.

On this basis, Meta has also designed a "Sentinel" security Agent specifically to monitor the data flow in and out of Muse, which will directly intercept and prompt user authorization if it detects anomalies or high-risk operations.

Metaverse, Smart Glasses, and Agents: Three Routes Converging

Zuckerberg revealed in the interview that Muse will be fully integrated into the Ray-Ban series of smart glasses and will upgrade existing interaction methods.

Currently, the glasses provide a "single-turn dialogue" experience—speak a sentence, get a reply, and that's it. After integrating Muse, the glasses become the front-end entry point for the Agent: users speak, and the back-end Muse continuously works in the Secure VM, completing tasks and providing feedback.

Regarding the product roadmap for the glasses, he described a clear hardware hierarchy:

  • Pure audio glasses without cameras: already equipped with Muse, capable of handling calls, music, and voice tasks
  • Meta Ray-Ban with a small display: already released, providing basic visual feedback
  • Full field-of-view holographic AR prototype: already released, which Zuckerberg said "will be very exciting"

He stated that Meta's long-term investment in glasses positions the company favorably after the AI Agent matures. Meta is introducing Muse and more AI features into the glasses products, while the metaverse technology, which previously emphasized "presence," continues to advance, though more resources are currently shifting towards Muse and the AI functions of smart glasses.

Regarding virtual avatars, Zuckerberg mentioned that previously, Meta needed room-level equipment, multi-angle scanning, and enterprise-level GPU environments to generate high-quality realistic avatars; now, users can complete the setup with just a few photos, and related capabilities can run on a VR headset.

Looking ahead to 2030, Zuckerberg stated that the core vision of the metaverse has always been to merge the physical and digital worlds. He gave an example that in the future, people could participate in activities with friends who join through holographic images; in work scenarios, humans and multiple Agents could also participate in group chats or meetings, with Agents appearing in holographic form or other embodied forms.

"The Scariest Moment": The Mistake of Llama 4

Not all bets have gone smoothly. Zuckerberg rarely spoke directly about the failure of Llama 4 in the interview.

After Llama 4, that was the scariest moment… I thought we were on the right track, but we weren't. It was quite a significant negative surprise.

He attributed the issue to a fundamental error in team structure:

I built the team in the way of Instagram's recommendation system or advertising system—hundreds or thousands of people working in parallel. But training language models requires a small team that collaborates closely, treating it as a collective scientific project. Every position is extremely valuable.

As a result, Meta underwent a complete reorganization, widely bringing in top talent from the industry and establishing the Meta Super Intelligence Lab (MSL). Zuckerberg stated that the next generation of models is about to be released, but it will not be announced at the Connect conference.

Computing Power Path: Brute Force AGI, 5-Gigawatt Project Under Construction

When discussing the technical path to AGI, Zuckerberg provided a straightforward judgment.

I'm not sure what fundamental architectural breakthroughs are still needed… I think we roughly know the formula. If you can build a sufficiently large supercomputer cluster, you can brute force your way there.

Currently, Meta's computing power expansion route: A cluster of over 1 gigawatt in Ohio is basically in use for training the next generation of models; a 5-gigawatt cluster in Louisiana is under construction.

He stated, "When you have multi-gigawatt clusters for training, you basically get something close to AGI or even superintelligence."

However, he also added that the current computing power-driven route does not mean that architectural research is unimportant—

The human brain consumes about 10 watts, while our systems may be a million times less efficient. Only by combining large-scale computing power with architectural breakthroughs can we truly lead.

Alignment is Not a Burden, But a Problem That Products Must Solve

Zuckerberg's attitude towards AI safety is very pragmatic—he does not see it as regulatory pressure but as a prerequisite for product success.

If you ask Muse to do something and it does the opposite, who would still use it? We need to ensure that the model not only understands your specific instructions but also comprehends your intentions and values.

"If Muse is to reach a billion users, we must solve the alignment problem or make significant progress on it," he said.

Regarding safety boundaries during the training process, he used an analogy: "It's like parents setting rules for their children—if it 'completes' a programming problem by modifying system configurations, I need to tell it: No, I want you to learn the problem-solving method, not to find a shortcut around it."

Zuckerberg Talks Muse: AI Agent Explosion,

The full interview is as follows: >

Major Bets

Host: This is something that billions of people will want to use. How does it feel for you to be building this? Is "immortality" a possibility? Well, if you take me into your mind, fast forward to 2030, what will that look like?

This is Mark Zuckerberg. 22 years ago, he built a social network that connected billions of people and changed the world forever. And now, he has decided to build something even grander. To this end, he is making major bets in the fields of the metaverse, smart glasses, and artificial intelligence. For years, skeptics have believed these are three separate bets that cannot succeed, but they have overlooked the bigger picture—because right now, these bets are converging to create a brand new super intelligence.

Today, we will present all of this, and I will ask Mark some questions he has never been asked before, listening to his vision for the future, allowing you to get ahead and build the next big thing.

Host: Thank you very much for coming to the show.

Mark: Thank you, I'm glad to be here.

Host: For this conversation, I watched every interview you've done.

Mark: Wow, that's more than I've watched myself.

Host: It was fun and very rewarding. Two things left a deep impression on me: first, your love for "building"; I feel you are one of the top builders; second, your ability to make bets—daring to make huge wagers. I feel that this week, all these bets are converging, so let's start from there.

Mark: Okay. We've been doing these things for a long time. In terms of AI, as a company, we've been working on it almost since our inception—the first version of the news feed was in some ways a machine learning product. Then about 15 years ago, we created the AI research lab.

But now we have entered a new phase—about a year ago, we launched the Meta Super Intelligence Lab. This is a pretty thorough research reboot, bringing in a lot of great talent from across the industry, which is exciting. Currently, we see the models getting better and better, and the next generation of models is about to be released, but it won't be announced at the Connect conference. Additionally, we have the Muse personal assistant, which has received very positive feedback so far.

When building these things, you are not really sure what the outcome will be. We like it ourselves. Earlier this year, I was at home piecing together Open Claw in my own way, feeling how this thing works and how to turn it into a magical experience that anyone can use.

Basically, when we—myself, Nat, and Alex—sat together and realized this was a magical experience, and if we could turn it into a plug-and-play version that ordinary people who don't understand technology, don't want to set up a Mac Mini themselves, don't want to mess with the terminal, and don't want to debug when something goes wrong could use, I felt this would be something billions of people would want to use.

Since then, we've been working towards this goal: specifically tuning the models for it, not only building the assistant itself and the operating framework but also building the technology that provides an independent computing environment for each assistant—we built a whole set of Muse secure virtual machines for this.

Internally, we always felt this was special, and the internal team liked it, but you never know how people will react after the product is released. Occasionally, there are cases where it becomes a hit right out of the gate, but most of the time you receive some positive feedback and need to iterate on a few things to really make the product "click." But this time, it clicked right away. Seeing all this happen is really exciting—within just two weeks, the number of users has already reached millions, which is quite rare. We only encounter such situations every few years, but this is definitely one of the most enjoyable moments for a startup.

Host: Your ability to stay in the game and keep trying is very admirable. And hitting so many home runs is really impressive. You had an interview with Joe Rogan about three years ago, where you mentioned that one day you could wear glasses, and the AI assistant would appear. So I feel that Muse already has a physical representation, which is very clever because that feeling is something that is bound to happen.

Mark: I think it just makes it seem friendlier and cuter. I believe too many people describe AI as something frightening, but AI should just be useful and fun. That physical representation—a designer made that character early in the project, and for some reason they always wanted to iterate, but the first version was the best. Later, someone said, "Oh, it has to be blue because it's Meta." I said, "No, I think this representation is right; you nailed it the first time." That's just how it is; it's fun.

What is the difference between personal AI and general AI?

Host: Great detail. If you want the model to perform exceptionally well in personal matters, rather than just being a general intelligence model, how would the training approach differ?

Mark: I think the core right now is to make the model an excellent "agent." Over the past year, the biggest trend has been programming agents, which contains two core ideas: programming expertise and the general ability to be an excellent agent.

Our strategy is to prioritize making the agent itself excellent, rather than specializing in programming capabilities.

Programming capability is important because even if users don't think they are writing code, Muse is writing code in the background to accomplish various tasks for you. But we believe the agent should first and foremost be an outstanding agent, with programming capability serving that goal.

Additionally, when building personal assistants, some things are more important than building enterprise software products. For example, if you are building an enterprise programming tool, the model doesn't need to have any concept of "information disclosure scale"—like what information should be shared and what shouldn't. But for Muse, this is very important.

You need to train this capability into the model, just like training any other capability. You will tell it a lot of things and then hope it helps you achieve your goals. For instance, if you want it to help you book a restaurant, you are looking for a suitable restaurant, but you might have some allergies, or you might be pregnant, and you may not want to disclose this to the restaurant. But Muse will know this information, and you want it to complete the task while disclosing as little information as possible. This is a specific skill, essentially basic social common sense for humans.

But most companies that only focus on programming agents have not trained this capability in. We can do this because we are doing full stack—we are not taking a ready-made model from someone else and just framing it; we are training the entire model from scratch, specifically to serve these capabilities.

The model certainly needs to have broad general intelligence, but it also possesses these specific capabilities. Then, on this basis, you build the entire assistant and all the details around it: memory, operating framework, and how it "beats"—it wakes up periodically to check, "Okay, these are the goals I understand about you; is there anything I can push forward now?"

We also have a dedicated team that focuses on refining the real-time animation effects of the virtual representation, because it's not just the default representation—you can customize any representation, and then it moves naturally, which is fantastic. We are also launching a voice mode, where you can have real-time voice conversations, and your assistant is right there with you. I think all these details stem from us doing full stack—the model and product are developed together.

Host: Another major thing is the virtual machine. Can you explain why it is important and what it unlocks?

Why does Meta provide each AI assistant with an independent computer?

Mark: Basically, for an agent to do things for you, it needs a place to store your information. We believe this information should not be mixed in a resource pool with everyone else's information.

You can think of it this way: your assistant will know a lot of sensitive information about you. Many people initially encounter agents like OpenCloud and set up a Mac Mini at home themselves. So we thought, many people do not want to buy a Mac Mini or set it up themselves. So what kind of experience can best approximate this effect? The answer is: you just need to download an app, register, and you will get a computer dedicated to you for your assistant to work, store data, and build a secure model around it. This is akin to having your own computer under your desk—even we at Meta cannot access it.

For example, the Muse confidential virtual machine is a feature we are developing, and even we cannot see the contents of your virtual machine.

We have also built a lot around this, such as secure credential storage—when your assistant needs to handle information like passwords, it itself does not need to see these passwords; it just needs to "insert" credentials when you ask it to log into a service, and only do so when you explicitly request it. The system should be designed so that this information is not freely accessible, because accidents can always happen; someone might try to hack in, or the system might have issues.

So you need to ensure that the agent cannot access this data, and neither can Meta. Providing each assistant with an independent computer and maximizing security is the fundamental basis of this technology—empowering Muse with the capabilities needed to help users achieve their goals while ensuring privacy and security, making it a world-class, industry-leading product in this field.

Host: Does this mean that, like WhatsApp, the data on the virtual machine that Muse connects to is encrypted? How should people understand the actual storage of data in the virtual machine?

Mark: We basically built two versions. The Muse Secure Virtual Machine contains various privacy features, including the entire Sentinel agent architecture we built. You have a regular Muse assistant performing tasks for you, and at the same time, there is a security agent we call Sentinel, specifically monitoring the data flow in and out of your Muse.

If external content attempts to compromise security, Sentinel will directly cut it off, preventing it from happening. If it believes your Muse is about to take an action that requires your intervention, it will override Muse's operation and trigger an alert for someone to confirm permissions—like "Do you want Muse to do this?"

This entire system, combined with secure credential storage and multi-layered defense, constitutes the Muse Secure Virtual Machine.

We are also developing another project. Nat and I specifically recruited Moxie Marlinspike—he is the one who collaborated with us to achieve end-to-end encryption for WhatsApp—to design the Muse Confidential Virtual Machine. The core idea is to give you a dedicated encryption key on top of the secure virtual machine, so that even Meta cannot access the contents within.

This version is more challenging to implement because if Meta cannot access the internal workings of the virtual machine, debugging and ensuring the system runs smoothly becomes much more difficult. So we took some time, but it will be released soon. This will essentially meet the security standards that people are already familiar with on WhatsApp and our other most secure products.

Host: Is the advantage of doing this just to make people feel psychologically safer, or are there actual benefits?

Mark: I think security is important in itself. Our goal is to approximate having a local machine under your table. What does that local machine give you? It means no company can access it.

So, assuming Meta wants to provide you with this service, how can we offer you the same level of privacy and security so that no company—whether it's Meta, anyone trying to invade us, or in some countries, a local government you don't trust—can access it? Because we can't access it either, as we have no access rights.

I think this is very important, and it's one of the key reasons people trust WhatsApp. There is real value in privacy, security, and trust.

If you were to have an assistant who knows everything about you—I guess almost all of us would have such an assistant—fast forward five years, everyone will have an assistant that understands your goals and everything about you and can help you get things done. In this case, being industry-leading in privacy and security is very important. We wanted to do this from the very beginning.

How to Shape the Personality of AI?

Host: You have an interesting point—you studied psychology in college.

Mark: Well, I was only there for a short time, two years, but I feel it influenced a lot of what I built later.

Host: When we look at models, we often say that a certain model is "very smart." But just like when we choose friends, it's definitely because they are smart, but also because we like their energy and how we get along. How do you consider shaping the personality of the model?

Mark: I think an ideal model should have enough adaptability to fit different people's styles. I feel many industry professionals have gone in the wrong direction on this—many other labs focus on "how to design personality correctly," but I never thought of personality as a fixed thing. This is also one of the reasons I strongly believe in open source and that people can customize, and why we designed Muse to be a highly personalized product.

You can customize and personalize Muse—not just its appearance and voice; the first thing it asks you when you sign up is "What do you want my basic personality to be like?" and you can modify it at any time.

We strive to make the model highly steerable, allowing you to define how you want to interact. This is a crucial part of making it a great personal assistant—this adaptability around personality is key.

Host: What style is your own Muse?

Mark: I have it direct and efficiency-focused. It's quite interesting. Earlier versions were very sarcastic and humorous, but this version is more straightforward. My assistant uses the default Muse appearance, but I dressed him in a toga and gave him a voice that's deep to the point of being comical; interacting with him is fun.

Host: I think to have a sense of humor, you must be really smart. Many people don't realize that comedians are among the smartest people in society—they need to be quick-witted and clever. You definitely have that trait. I've watched all your interviews, and you've always performed well.

Mark's Surprising Predictions About AI and the Metaverse

Host: In your interview with Theo, you talked a lot about the next frontier of technology and where all this is ultimately heading. The AI field has gone through several "winters," where people thought there wouldn't be breakthroughs. The metaverse has also gone through several such periods, where everyone thought it was an unfulfillable bet. I tried out the new holographic avatar feature yesterday, and it's very cool. In the interview with Lex, it felt like it would take 11 hours to record your face, but now it only takes 3 minutes. How did we get to this point today?

Mark: In the overall development of the metaverse, when we founded Reality Labs, we always believed that eventually, there would be normal-looking glasses, and over time, they would provide immersive presence while also being excellent AI devices—because glasses are the only form that can see what you see and hear what you hear, communicating with you all day long and ultimately displaying images.

But 10 to 15 years ago, I assumed we would first achieve holographic technology, then have highly developed AI. However, the path of technological evolution is interesting—we actually had AI and personal superintelligence first, and then the technology to make holographic technology sufficiently widespread and affordable came along. This was something I didn't foresee, but I'm glad we're working on both directions.

Our significant investment in glasses puts us in a very advantageous position as AI assistants become ready. Many of the releases at the Connect conference were about bringing Muse and a lot of AI features into glasses, which I think users will love. This is a big deal.

Regarding presence, we are still making progress, but the pace is relatively slow because most of our energy has shifted to building Muse and AI features for glasses. However, we have a long-standing project focused on real-time high-fidelity virtual avatars.

As you mentioned, three or four years ago, you needed an entire scanning room to capture a person from various angles, and you needed enterprise-grade GPUs to render it, which was very cumbersome. The demonstration we did for Lex's podcast was that kind of setup. And now we can basically run it in a pair of VR glasses—this is the first form of glasses that can deliver such an amazing VR experience, rather than a bulky headset. You can create your virtual avatar with just a few photos, and the progress is astonishing.

Host: And it can also drive expressions based on voice. In the demonstration, it made me laugh, let me interact, and then understood how my face moved with the audio track. You mentioned in that podcast that some people who are usually more reserved with their expressions actually want richer expressions in the virtual world. What do you think about people distinguishing between their "virtual self" and "real self"?

Mark: I think we are still in the early stages of understanding this sociologically and psychologically. I feel that people's self-perception and the image they want to project often differ somewhat from their actual selves. Since the inception of social networks, people have been carefully selecting their avatars. We see a similar phenomenon with Muse's virtual avatars. It's not so much about "curating" oneself, but about "curating" the "person" you want to communicate with.

I believe that when you provide people with the ability to express themselves, you want it to genuinely capture them, while also being a form of expression in itself, not just a pure mirror reflection—it is both communication and expression. We hope to build something that balances both. This is always an iterative loop: seeing how people use it and then improving it. After years of work, we are now truly at the starting line—able to incorporate quite high-quality realistic avatars into products that can be used on phones and in VR. I am very excited to see the results.

What Will 2030 Look Like?

Host: Okay, take me into your mind, fast forward to 2030. If all goes well, what will holographic technology look like? How will holograms, Muse, and glasses converge?

Mark: My understanding of the metaverse vision has always been about effectively merging the physical world with the digital world. The basic idea is: we have this beautiful physical world, and we also have an amazing digital world—the vast content accumulated on the internet over the past 20 to 30 years is breathtaking. But the way we access it is either sitting at a desk or through a small screen in our pocket, which is fundamentally very limiting.

I think the ideal version of this is a seamless integration of the physical and digital worlds. You can think of it this way: right now, the two of us are here. In a future version, one of us might be a holographic projection, but you still feel the genuine presence of each other, which is completely different from a video call.

The core of virtual reality is to convey this sense of presence—making you truly feel like you are in the same room with others or in another place. You can achieve this with holographic technology and mix it in various ways. For example, I could play a game of poker with friends, with some present and others joining via holographic avatars, and the poker table itself could also be holographic, allowing those who are not present to be integrated into the game.

At the same time, AI can also be given a physical presence in this scenario. This makes a lot of sense in work contexts—I am currently using various programming agents to build things. Imagine this: you have a group chat channel with several people and several agents, and you assign tasks to the agents. But sometimes, everyone gathers for a meeting, and the agents should also be present. How do they appear? It's simple; there are a few extra seats on the couch, and they appear as holographic avatars. Or you could use Muse's cute little character, or a dragon, or any quirky image you create.

I guess this will feel quite natural in the future.

Host: Interesting. You mentioned in previous interviews that the tech industry often forgets about the "fun" aspect. I think having a physical assistant present would also make it feel more real—it’s like truly outsourcing work. When you see Muse typing, it feels like something is really happening. This is done well—you can see what's happening in the browser. So do you think there's a possibility of a scenario where you wear glasses and control your computer, letting Muse help you do things on the computer?

Mark: Oh, definitely. VR can already do that. You can sit anywhere, even in a café, and open your workstation with six monitors to write code; everything is there.

On the glasses side, the most popular model currently doesn't have a display, which makes it more affordable for more people to use, while we are still working to make displays into the most compact form. But we have released a display version of Meta Ray-Ban, which is very popular; it's a small display. We also released a prototype version of full-field holographic AR, which I think will be very exciting.

So the entire product line looks like this: from pure audio glasses—no cameras, looking just like regular glasses, but with Muse inside, allowing for various audio tools, listening to music, making calls—to higher-end versions, all of it.

Host: I'm wondering, with those audio glasses, can you talk to Muse while letting your home computer do things?

Mark: Yes, absolutely. We just launched this feature at the Connect conference. Now all glasses are connected to Meta AI, creating a "single-turn" experience------you send a prompt, it replies, and that's it. But with Muse, we basically want to upgrade all glasses to Muse. First, you no longer need to say "Hey Meta," you can name it anything, which is part of the fun. Then you just talk to it directly, it connects to your Muse, and your Muse processes tasks in your secure virtual machine, helping you get things done.

Host: That's amazing!

Founder Mindset

Host: Alright, we are here now, Muse is progressing smoothly, and the glasses are doing well, but about a year ago, many people were asking, "What happened to the Super Intelligence Lab?" At that moment, how did you feel inside? What was it like when things weren't going well, but you still saw the long-term vision?

Mark: The real problem was with the Llama project and Llama 4.

Llama 1 was quite an interesting model; it pioneered the entire open-source AI movement, and we are very proud of it. Llama 2 achieved scalability, Llama 3 was a great model that was almost at the cutting edge at the time. Then with Llama 4, we basically deviated from the development trajectory we should have followed.

Whenever things don't go in the direction I expect, I spend a lot of time thinking: Why is this happening? What do we need to change to do better? This time, my reflection was: I got the entire team structure wrong.

I modeled it like the way we do machine learning work for Instagram feeds or advertising systems------with hundreds or even thousands of people working in parallel on many things. But for building language models, what you really need is an extremely tight-knit small team, treating it as a collective scientific project. You don't need many people, but this means that every position in the team is extremely valuable.

So we recruited the best people from all over Meta, while also bringing in many talented individuals from the industry to form a brand new team------the Meta Super Intelligence Lab.

From my perspective, when MSL launched, I knew it would take time to restart, rebuild infrastructure, and train the next generation of models. But I knew we had assembled an excellent team, and if the team could gel and operate well, the results would be good.

For me, the most thrilling moment was actually after the release of Llama 4------I thought we were on this track, only to find out we weren't. That was quite a significant negative surprise. I think, as an entrepreneur, you are always tested in these moments because inevitably, not everything will go smoothly. What truly determines the trajectory of development is: how you find a way forward when things don't go the way you hope.**

Host: I guess it's the opposite as well------when something far exceeds your expectations, like the launch of Muse, how do you ensure you seize that opportunity?

Mark: Absolutely. Now the entire company is invested------initially, it was just a small team building the product, but now everyone realizes that it is really ready to take the big stage. The whole company is thinking about how to scale this thing up, how to let hundreds of millions of people experience it.

From optimizing all infrastructure to ensure everything runs smoothly, squeezing every bit of computing power from existing GPUs, to various product teams embedding Muse in different ways------like in glasses. Seeing everyone working together to ensure Muse can scale out smoothly feels fantastic.

How Mark Writes and Communicates Vision

Host: As a founder, I feel that this is the moment you long for most------everything comes together. How do you communicate your vision to the company? Given that you are pushing so many things forward at the same time, I feel you write a lot. What is your process?

Mark: Writing is very helpful for me, both to refine my thoughts and to communicate externally. This summer, I wrote a long article called "The Future Belongs to Everyone," about 15 pages long, which helped me systematically clarify my philosophical stance on various important social issues related to AI: what I think is good, how interactions with the government should be, how to prevent various harms that people are concerned about, how to make data centers a wealth for the community, truly creating jobs instead of destroying them, how to maintain national security, and how to effectively alleviate the risks that people worry about------whether it's hacking or biosecurity issues.

This is a very complex matter, taking a long time, communicating with many people, and going through many rounds of revisions, with many people internally participating in discussions and debates. But for me, it was a very valuable process. In the end, we had something like this------15 pages, "This is what we believe." Then I distilled it into a one-page version, published as a column. We also made a short video because I feel that to reach many people, you often don't need to throw out a whole theoretical argument, but rather condense it into: what your values are, what you believe, and how you communicate that.

There is no one-size-fits-all way to communicate these aspects to the company or the world. Different times require different approaches, and different groups of people, some naturally resonate with what you are doing, while others are more concerned and need you to guide them, requiring extra effort to explain why this is valuable. I think that is part of running a company and being a founder------you are not repeating the same thing over and over; each situation is slightly different, facing new challenges, which is part of what makes it interesting.

What Drives Mark to Build?

Host: I feel that for you, this founder mindset extends beyond the company------whether it's your farm or learning a new skill.

Mark: I just love to build things.

Host: What does "building" feel like for you?

Mark: I feel it is an inner need. Different people have different ways of self-expression. If you are a writer, you feel you must write. Some people have a need for recognition. But I just need to build things. If I'm not exercising my creativity or building something, I become irritable. That’s not pleasant for those around me.

Host: Is learning to be a great skier the same skill as learning to build a product?

Mark: To some extent, yes, the process of learning new things is quite similar. In my life, I have deliberately challenged myself with things I am not good at. For example, I have always been very bad at learning languages. This is actually why I started learning Latin------I just couldn't learn French and Spanish in class, so I thought, well, Latin doesn’t require speaking, just translating, like math.

Later, when I started running a company, I set annual challenges for myself, one year was to learn Mandarin. Mandarin is really hard, especially the tones. Of course, there are good reasons to learn it------Priscilla's grandmother only speaks Mandarin, so if I want to communicate with her, I need to learn. But the biggest motivation is actually the challenge itself.

All these things, you can only do them; there is no shortcut to "understand." You can only invest time, and then it slowly seeps into your brain. Learning martial arts, learning to fly a helicopter, is the same------these things are actually hard to "understand" rationally; they require practical accumulation.

Building products is partly like that. Programming can be thought of in theoretical ways, but the intuition for building products can only be cultivated through repeated practical experience. The only question is what you enjoy------because I believe not everyone has the same strong need to build as I do; most people have some degree of need, and the key is to find out what that is and then invest time into it to become truly excellent at it.

Honestly, I feel this is not just a "want," but a deep psychological need or drive. Aligning this drive with something and giving yourself time for those experiences to slowly seep and settle is crucial. This is also what I try to teach my kids------to help them find what they are truly interested in, but if they hit a bottleneck, to give them a push.

One of my daughters loves to create music, but she just can't stand taking piano lessons. I told her, "You don't need to become a piano master, but if you want to create music, you need to have an intuitive grasp of music theory and how it works. And once you learn piano, you can immediately pick up the guitar." I think, whether as a child needing a parent to give a push, or as an adult, having the discipline to sit down and let these things slowly settle in the brain is key.

The Golden Age of Builders

Host: I feel that actually everyone wants to build. I think Generation Z is not as criticized by the outside world------low motivation and so on. I feel people are just looking for the spark that ignites their building power. You talked a lot about this in your speech at Harvard.

Mark: Yes. When I say "building," I mainly refer to these kinds of products------software, hardware, etc. But I agree, I think everyone has some kind of creative drive. However, some people have other stronger personality traits that overshadow this, such as service-driven------those who become doctors or nurses, their core drive is "I just want to take care of others."

I've heard a story, possibly from Priscilla's experience in medical school: on the first day of class, someone stood up and asked, "How many of you have memories from childhood of seeing someone and thinking, 'I really want to take care of that person'?" As a result, everyone raised their hands. For me, my version is: I have many childhood memories of "I want to make this thing better."

Not everyone has the same drive, but everyone has something they want to do. I agree that with previous technologies, starting out has been difficult for many. This is also what excites me most about personal super intelligence, Muse, and various AI assistants------I believe this is the first time in history that people can really get started quickly. You have a vague idea, and AI can help you outline it, and then you can refine it, like sculpting a sculpture, without needing to understand everything before you start.

I think this is very powerful, and many people will find what they want to create or push forward. This will help people feel a broader sense of agency.

How Far Are We from Conquering All Diseases?

Host: Another project you have outside of Meta is conquering all diseases. I'm curious, first, how close are we to that goal?

Mark: Much closer than before.

Host: Realistically, how close do you think we are?

Mark: When we initially launched this project, the goal was to help the scientific community conquer all diseases by the end of this century. We were never going to do it ourselves; our theory is: all major scientific advancements are preceded by a new tool that can measure and understand certain things. For example, the invention of the microscope helped us understand bacteria; with the telescope, we understood the universe.

Some of these tools have become platforms------for example, the first person to invent a vaccine, after that, people could use this method to treat many diseases.

However, historically, the way scientific funding has been distributed has been widely decentralized, allowing individuals to explore independently, and there hasn't been much funding actually used to build these large tools. What we are trying to do at Biohub is to design several new tools that empower people to "see" biology in new ways, thereby helping the scientific community accelerate progress.

Initially, we thought that "by the end of this century" was a very achievable goal; many biologists at the time felt it was impossible. But now I feel that the end of this century is too far away.

The progress of AI, combined with the virtual cell models we are working on—the basic idea is to conduct experiments not on real living cells but to use AI models to simulate proteins, then simulate cells, and further simulate virtual immune systems or entire organisms, even entire humans. This will allow scientists to run a large number of experiments and simulate what happens under different circumstances, such as what reactions occur in the body when a person takes a certain medication.

I don't want to give a precise year, but I guess it will be much earlier than the end of this century.

Host: The goal of "conquering all diseases" feels very deliberate, rather than "immortality." Is immortality a possibility?

Mark: This is not really my area of expertise. I believe the two are different things. "Immortality" is more about extending life—even if you don't get sick, the human body has a natural life expectancy, which is another issue that needs to be studied separately. Some people might consider this a "disease," which is a reasonable perspective, but I think people also need to work on conquering those diseases that will still make you sick even after you solve the lifespan issue.

The part of the problem we are choosing to tackle does not mean you will never catch a cold. Our statement is "cure, prevent, or manage": some diseases can be completely cured; some I believe we will be able to prevent in the future; and some, perhaps you will still get sick, but what would have caused real harm or even taken your life, you can now manage as a chronic condition without significantly affecting your quality of life.

The goal is to keep the human body in a balanced state—not that we will never come into contact with pathogens, but that one day we will be able to cure, prevent, and manage all diseases.

What is the most frequent thought in Mark's mind?

Host: AI has acted as a catalyst in many different breakthroughs, including personal intelligence. What do you think about most right now? What thoughts come to your mind most during this process?

Mark: This may change every week.

Right now, I am very focused on Muse. Every time we release something, we push forward quickly, then engage with the real world to understand people's feedback, and figure out what to do next—this phase is always exciting. We are currently in this phase, collecting a lot of information to understand what people want. The good news is that people really like it, so we are working hard to get as many people to use it as possible.

There is still a lot of work to be done on the model side; the Muse Spark model has made significant progress, and we hope to continue pushing forward to create a world-leading model.

What breakthroughs are needed next?

Host: What breakthroughs are still needed?

Mark: When researching this kind of thing, you may not be able to predict in advance. We have some hunches about certain directions, and much of the work over the past year has actually been about expanding infrastructure.

About a year and a half ago, more people would say that to achieve superintelligence, certain fundamental architectural breakthroughs were needed. But I am not so sure that is the case now. I think we have a pretty good understanding of the "recipe"—if you can build a sufficiently large supercomputer cluster, you can reach there through brute force computation.

We are building that cluster in Ohio, which exceeds 1 terawatt, and it is basically online; we are using it to train the next generation of models. Then there is the 5 terawatt cluster in Louisiana, which will be put into use soon. I estimate that when you have a multi-terawatt scale cluster doing training, you can basically get close to AGI, and even superintelligence above AGI.

But that doesn't mean this is the best way. The human brain only needs 10 watts to operate, while the computing systems we are building are about a million times less efficient than the human brain. So I do believe there is room for improvement in architecture, and we are also exploring—if you combine massive computing power with architectural improvements, you can really build something that leads the world. But for now, scaling itself can take us a long way.

Host: Indeed, scaling is the most reliable path. In research, you have to hope for major breakthroughs, but scaling can yield results quickly.

Mark: That's why large companies choose this path. There may be a much cheaper way, but we don't know it yet. And this is something so valuable to the world that even if it costs hundreds of billions of dollars, as long as the probability of success is high enough, it is still worth doing—ensuring that even if a cheaper breakthrough is not found in the end, you still have a clear path to achieve it. Of course, if you find a cheaper way, that would be even better.

So I don't want to say this issue is "solved," because at every stage of scaling infrastructure, you will encounter various new engineering challenges that need debugging and resolution. Research itself is advanced in this iterative way.

How to Win the Battle for AI Safety

Host: One of the most important things right now is to focus on alignment, making the models trustworthy.

Mark: Yes. There is a debate in the industry about alignment: will these AI labs naturally do it? My view is: of course they will. If you tell Muse to do something and it does the exact opposite, I don't know how many people would still want to use it. We must ensure that the model not only understands what you specifically asked but also understands your intentions and values, so it doesn't do things in a way that you are not satisfied with or produce negative effects that you absolutely do not want. This deep understanding is essentially alignment.

So my view is that for Muse to succeed and reach billions of people, we must make tangible progress on alignment, not just talk about it.

Host: Is alignment achieved through users saying "that's not what I want," or through your backend discovering it?

Mark: Both. User feedback is there, but I think it is becoming increasingly clear that you need to do alignment during the training phase, not just after deployment. Now that the model is smart enough, if you don't do it well during the training phase, most of the safety issues and incidents we see in other labs occur during the training process, not after deployment to users.

You need to set up a very good "curriculum" for it, just like parents teach children, setting clear boundaries. If it does the wrong thing, it needs to learn "no, you should really solve this programming problem, not bypass it by modifying some system configuration to get a reward"—I want you to learn this method through the actual problem-solving process; that is the meaning of training.

This is largely about establishing good safety mechanisms and clear boundaries. These are all natural tasks that the industry must undertake, and we are investing a lot of time; the situation changes every day.

Host: I remember you mentioned that the issue of AI safety—seems to have suddenly become a focus in the last two weeks, but actually, there wasn't a single event that triggered this timing. It sounds like you mean we actually spent a few extra months on internal training.

Mark: For Muse, we know this product needs to pay very close attention to privacy and safety. We have an early version of the model that we believe can further train some behaviors regarding "information disclosure scale"—just like we talked about before. So we took the time to do that. We also spent time making the virtual machine more secure. This took a few extra months.

I don't quite agree with some other labs' statements—"we are suffering greatly to slow down." For me, that is the right thing for Meta and Muse users. We want to make the product good; if the product is not good, it is bad for users and bad for us—we don't want to release something that leaves a bad first impression, and then everyone loses interest in using it.

I believe all labs have a very strong intrinsic motivation to get this right. Aligning the incentive mechanisms with the goal of "extremely valuable and safe" is key.

Host: Thank you very much for taking the time during this important week.

Mark: Thank you, thank you.

Join ChainCatcher Official
Telegram Feed: @chaincatcher
X (Twitter): @ChainCatcher_
warnning Risk warning
app_icon
ChainCatcher Building the Web3 world with innovations.