The benefits of medical AI assistance vary based on user expertise

A one-size-fits-all approach likely isn’t the best strategy when designing artificial intelligence systems that assist users in disease diagnosis.

A new study by researchers at MIT and elsewhere found that, while AI assistance generally improved the accuracy of non-experts and clinicians in diagnosing skin diseases, AI explainability methods had different impacts depending on the users’ knowledge level. 

Explainable AI methods help users know when to trust a model’s predictions by describing or validating the model’s decision-making. For instance, a model might use a heat map to highlight image regions that were most important in its diagnosis or a large language model (LLM) to explain the prediction in plain language.

In this study, researchers tested non-experts and primary care providers in skin disease diagnosis, with and without the help of different explainable AI systems. 

They found that non-experts’ diagnostic accuracy improved, but it was largely due to deference to the AI system. Non-experts trusted LLM-based explanations whether they were right or wrong, and found explanations more convincing when they were vague or generic.

By contrast, clinicians were not tripped up by incorrect AI assistance and performed best when given only a model’s prediction, with no accompanying explanation. 

“Good AI systems can improve performance in some health settings, but this has to be balanced carefully with algorithmic deference that can lead to more error. We know that both AI and explainability methods can engage automation bias in humans, and this anchoring effect is something that must be accounted for when we design AI systems,” says Marzyeh Ghassemi, an associate professor in MIT’s Department of Electrical Engineering and Computer Science (EECS), a member of the Institute for Medical Engineering and Science, and a principal investigator at the Laboratory for Information and Decision Systems and the Abdul Latif Jameel Clinic for Machine Learning in Health.

“These findings are important as patients increasingly turn to AI to help with their health care. Our findings show that those with the least medical knowledge are most likely to be led astray when explainable AI models give an erroneous output,” says Roxana Daneshjou, a co-author and assistant professor of biomedical data science and dermatology at Stanford University.

These results underscore the importance of building AI systems with users in mind and of developing explainability methods that encourage critical thinking rather than overreliance on the model, the researchers say.

“It’s getting obvious that we cannot just assume a good AI will solve all problems. We need to pay careful attention to the users who will be using the AI system, because the same explanation can help an expert and mislead a beginner. Often the people who could benefit most from AI are the ones most likely to be led astray by it, so how we present a recommendation matters as much as whether it’s correct,” says lead author Orson Xu, an assistant professor in the Department of Biomedical Informatics at Columbia University.

Ghassemi, Xu, and Daneshjou are joined on the paper by many authors, including MIT graduate student Haoran Zhang, undergraduate Reina Wang, and Luis Soenksen PhD ’20, a research affiliate at the Jameel Clinic, along with clinicians and researchers. A description of the work appears today in Nature Medicine.

Exploring explanations

Several FDA-approved AI interfaces are being used to help clinicians identify skin conditions in medical images, as a way to streamline early diagnosis. In addition to providing a prediction of whether disease is present in the image, these tools often use one of several methods that explain the model’s decision-making.

At the same time, non-experts can perform digital diagnosis on their own using AI-powered search engines that predict skin diseases based on user prompts. These systems often use LLMs to explain the model’s prediction in simpler terms.

The researchers explored the effects and potential benefits of these explainable AI tools on primary care physicians and non-experts in dermatological disease detection. They tested users by showing them medical images plus an AI prediction of skin disease, employing different explainable AI approaches. 

These approaches included: an AI prediction and confidence level with no explanation, a method that provides similar images to reinforce its prediction, a heat map-based approach that highlights important image regions, and an LLM that explains the model’s reasoning in plain language.

Non-experts were tasked with deciding whether an image of a skin mole was cancerous, with and without the help of explainable AI. Clinicians were given the more challenging task of providing a differential diagnosis of dermatological disease.

The researchers found that all explainable AI approaches improved the accuracy of non-experts, mostly because the tools helped users diagnose non-cancerous moles. 

In addition, when they employed a fairness-constrained model designed to combat bias against darker skin tones, the system significantly improved accuracy and reduced diagnostic disparities based on skin tone.

“But the reason non-expert users are better is because they are more reliant on the models. When the model is wrong, it hurts performance more than it helps performance when the model is right. We were just able to train very good AI models for this setting,” Ghassemi says.

This deference effect is largest with LLM explanations, and users were more confident about their wrong answers when aided by an LLM.

On the other hand, clinicians were resilient to incorrect AI explanations and, of all the explainability methods, LLMs boost their accuracy the least.

“It really comes down to how each group uses the explanation. A clinician already has a diagnosis in mind and checks the AI against their own training, so a bad explanation gets caught. Meanwhile, a non-expert can use that exact same explanation to form an opinion in the first place, so a plausible, confident-sounding rationale can pull them toward the wrong answer. The same tool ends up being an asset for one user and a liability for another,” Xu says.

Overcoming the deference effect

When the researchers dug deeper, they found that users who were most deferential to AI assistance were the worst performers on the task without the help of AI. 

They also found that the time at which users were presented with AI explanations influenced their behavior. If an explanation is given first, before the user can perform the diagnosis on their own, they tend to become more deferential to the model.

In addition, AI systems outperformed humans when the presentation of disease was subtle, but humans performed much better if there are atypical symptoms or unrelated features in an image.

Taken together, these results indicate that explainable AI can cause overreliance on models and lead users to blindly follow AI recommendations even when they are wrong. 

Rather than using LLMs to generate more detailed explanations, it might be more effective to force users to give a diagnostic hypothesis first, then provide an AI-based suggestion to highlight other possible conditions for consideration. 

“We really want AI to improve creativity and either upskill or fill in gaps where users are missing subtle presentations. Otherwise, we risk engaging automation bias and then, when the model is wrong, users can’t recover,” Ghassemi says. 

This research was funded, in part, by the National Science Foundation, Schmidt Sciences, the National Bureau of Economic Research, and Columbia University.

Alexander Rakhlin named director of the MIT Statistics and Data Science Center

Alexander “Sasha” Rakhlin PhD ’06, the Distinguished Professor in Data, Systems, and Society at the MIT Institute for Data, Systems, and Society (IDSS); and a professor of brain and cognitive sciences at MIT, has been named the next director of the MIT Statistics and Data Science Center (SDSC). 

Rakhlin succeeds Ankur Moitra, the Norbert Wiener Professor of Mathematics, associate director of the IDSS, and a faculty member in the MIT Department of Electrical Engineering and Computer Science (EECS) who has been SDSC director since 2021. Philippe Rigollet, the Cecil and Ida Green Distinguished Professor of Mathematics and a core faculty member in IDSS, also served as interim director in 2024-25.

“Sasha is one of the sharpest theoretical minds working in statistics and machine learning today, and also one of the most devoted mentors I know,” says Fotini Christia, the Ford International Professor of the Social Sciences and director of IDSS, which houses SDSC. “He has helped train an entire generation of interdisciplinary scholars through the Interdisciplinary Doctoral Program in Statistics (IDPS), while his own research keeps pushing the boundaries. The SDSC could not ask for a more fitting leader.”

Rakhlin is the inaugural holder of the Distinguished Professorship in Data, Systems, and Society, an endowed chair created in 2025 by the generosity and vision of IDSS professor Richard “Dick” Larson, an “MIT lifer” and pioneer in operations research, queueing theory, and system optimization.

“I am honored to take on this role,” says Rakhlin. “The strength of the Statistics and Data Science Center has always been its people — students, postdocs, and faculty from across MIT who bring sharply different perspectives to the most interesting problems of the day in statistics, machine learning, and AI. My goal is to support that community as it takes on the constantly evolving questions reshaping the field.”

Rakhlin has been connected to the Statistics and Data Science Center as a visiting professor since 2016, before formally joining MIT in 2018 in the Department of Brain and Cognitive Sciences and IDSS. As the initial chair of the Interdisciplinary PhD in Statistics program at the SDSC, Rakhlin has seen the successful defense of over 75 IDPS PhD students across a variety of departments at MIT, including IDSS’ own Social and Engineering Systems program.

“I have been fascinated by machine learning since my PhD work more than 20 years ago, drawn by its beautiful connections to statistics, probability, algorithms, optimization, and game theory,” says Rakhlin. “At the Statistics and Data Science Center, I work alongside colleagues who share this fascination and pursue these connections in many directions. The recent revolution in AI is extending this web into the sciences; it promises to accelerate discovery, and it raises new questions for statistics. Answering them demands a rigorous science of the tools themselves. As AI enters medicine, energy, and public life, its safety and security are, at their core, statistical and mathematical questions: quantifying uncertainty, providing guarantees, understanding failure, and resisting manipulation.”

As Rakhlin puts it, the SDSC is built for this moment. “Statistics is a shared language across MIT,” he adds. “Through the Interdisciplinary Doctoral Program in Statistics, the center connects students and faculty from economics and political science to physics and engineering. Collaborations in areas from biology to nuclear fusion have shown how statistical thinking accelerates science itself.” 

As director, one of his goals is to deepen these interdisciplinary connections. He hopes to help make SDSC the Institute’s home for the rigorous foundations of data science and AI, and a bridge to the scientific and societal questions where those foundations are most needed.

Rakhlin received his bachelor’s degrees in mathematics and computer science from Cornell University, and doctoral degree from MIT. He was a postdoc at the University of California at Berkeley in EECS before joining the University of Pennsylvania, where he was an associate professor in the Department of Statistics and co-director of the Penn Research in Machine Learning center.

A better way to turn 2D designs into 3D models for rapid prototyping

Engineers often use vision-language models to produce new designs, such as for airplane or automobile components. To simulate how those components will perform in realistic situations, they’ll use tried-and-true computer-aided design (CAD) software to generate 3D models of those designs, which they can put through virtual crash or durability tests. 

Researchers from MIT and elsewhere have now developed a system that can teach a vision-language model to automatically convert 2D designs into CAD programs that are much more accurate and functional compared to other approaches, while using only a fraction of the computation.

By improving the performance and efficiency of AI-driven CAD generation, this technique could streamline the rapid prototyping process and reduce costs. It could also help engineers identify beneficial design choices they might otherwise overlook. 

The system generates new data based on the model’s abilities as it attempts to convert a 2D image into a CAD program. The framework corrects the model’s failures and incorporates them into a dataset with its successful solutions. 

It uses these data to teach the model how to fix specific mistakes and tackle tricky problems it would struggle with on its own.

“We want engineers to be able to point our framework at an underperforming CAD model, set a compute budget, and let the system take over — turning the model’s own mistakes into better training data,” says lead author Giorgio Giannone, a research affiliate in the Design Computation and Digital Engineering (DeCoDE) Lab at MIT and a principal research scientist on the AI Innovation Team at Red Hat.

He is joined on the paper by Anna Claire Doris, a mechanical engineering graduate student at MIT; Amin Heyrani Nobari, an MIT postdoc; Kai Xu of RedHat; and co-senior authors Akash Srivastava, director of Core AI at IBM and a principal investigator at the MIT-IBM Computing Research Lab; and Faez Ahmed, associate professor of mechanical engineering at MIT, leader of the DeCoDE Lab, and a principal investigator at the MIT-IBM Computing Research Lab. The research was recently presented at the International Conference on Machine Learning.

“Nearly every physical product around us, from airplanes to appliances, begins its life as a CAD model. Industry teams are eager for AI that can help speed-up the creation of these designs, but today’s models often produce simple shapes inadequate for practice. What excites me about this work is that it gives many image-to-CAD-code models a way to improve themselves, learning from their own errors rather than waiting for more human-made data — and that brings trustworthy AI design tools much closer to everyday engineering,” says Ahmed.

Model-aware data

The researchers are working toward building vision-language models (VLMs) for CAD generation. These VLMs take a 2D image and some descriptive text, and output Python code that can be executed in a CAD software program to generate a 3D model of a physical object.

They studied the challenges of deploying existing VLMs for this task and determined the main bottleneck that limits their capabilities is the lack of diverse, high-quality CAD datasets to train them. 

To remedy this, they sought to create new data to teach a model how to perform CAD generation, using a process known as data augmentation.

In data augmentation, scientists typically create new data by randomly tweaking existing data to generate more samples, often by adjusting the color, size, and shape of objects in images. 

Instead, the MIT researchers built a data augmentation system called GIFT (which stands for Geometric Inference Feedback Tuning) that generates data designed to improve the performance of one VLM for a specific task.

GIFT develops an understanding of the model’s strengths and weaknesses by testing it. Then it uses this knowledge to generate data that could improve the model’s performance on the CAD generation problems it struggles to solve.

“We want to obtain data augmentation that is informed by the model itself,” Giannone says. 

Learning from mistakes

To do this, GIFT asks the model to generate code that solves a CAD generation problem multiple times in parallel. It checks the correctness of these guesses to understand how well the model can solve this problem.

“For a model, generating CAD query code that is almost correct is not that hard, but generating code that is perfectly correct and can be executed is much more challenging for a standard VLM,” Giannone says.

For guesses that are nearly correct, GIFT adjusts them to become successful solutions. It saves these “near-misses” and successful solutions in a new dataset that can teach the model how to overcome problems that would usually trip it up.

“If we sample the model 10 times and it generates 10 correct answers to the same problem, then there is not much for it to learn. We care about the in-between cases, where the model might only solve the problem 50 percent of the time,” he says.

Using these in-between cases allows GIFT to generate data augmentations that are both model-aware and task-aware. In addition, by incorporating multiple correct solutions to the same problem, the new data expand the model’s general knowledge of CAD code generation.

This automatic system does not require human intervention to correct the model’s mistakes.

GIFT creates data augmentations from a pre-trained VLM using a process known as inference-time scaling. This process allows a static model, which has already been trained, to generate better outputs without the high computational costs of retraining the entire model. 

Using inference-time scaling, the user can determine how much computation they want to use for GIFT, tailoring it to their time and budget constraints. 

GIFT outperformed several competing techniques, generating CAD programs that were more accurate while using only about 20 percent as much computation. The CAD models generated by VLMs using GIFT were better aligned with the shapes of ground-truth models.

“With GIFT, we started with geometry because with engineering problems, if the geometry of a 3D shape is not correct, nothing else will be correct, but there are many other aspects to consider,” Giannone says.

In the future, the researchers want to expand GIFT so the framework can teach models to generate CAD programs that improve the performance and manufacturability of 3D models. They also want to apply the system to larger models and more diverse CAD generation tasks.

This research was funded, in part, by the MIT-IBM Computing Research Lab. 

3 Questions: Neural transparency and the future of AI design

Millions of people are now designing their own personalized artificial intelligence companions, yet most have little idea how those creations will actually behave. In a new paper, MIT Media Lab Assistant Professor Pat Pataranutaporn and his graduate student researchers Anthony Baez and Sheer Karny introduce “neural transparency,” a tool that lets everyday users glimpse inside an AI’s neural network before their chatbot ever says a word. The work is being presented this week at the ACM Conference on Intelligent User Interfaces. 

In this interview, Pataranutaporn, who is the Asahi Broadcasting Corporation CD Professor of Media Arts and Sciences, explains what they found, why the stakes are higher than most users realize, and what genuinely transparent AI might look like in the future.

Q: Your paper introduces “neural transparency,” a way to let everyday users peek inside an AI’s neural networks before their chatbot ever says a word. Can you describe how that actually works, and why you focused on the design moment, rather than catching problems after a chatbot is already out in the wild?

A: Millions of people are now creating personalized AI chatbots and agents powered by large language models, turning them into collaborators, tutors, coaches, creative partners, and companions through simple text prompts. Yet most people have very little idea how those prompts will shape the AI’s behavior until they begin interacting with it. We wanted to change that.

“Neural transparency” means giving people something like a brain scan for AI. Not because AI has a human brain, but because its neural network contains internal patterns that can hint at how it may behave before it speaks. In this work, my students Anthony Baez, Sheer Karny, and I combined insights from the fields of human-AI interaction and mechanistic interpretability to make those hidden patterns accessible to everyday users.

The basic idea is simple. First, we choose behaviors we care about, such as empathy, honesty, toxicity, hallucination, or sycophancy. Then, we compare the model’s internal activations when it is prompted to exhibit one trait versus its opposite. That difference becomes a kind of “behavior direction” inside the model. When a user writes a custom system prompt — the instructions that shape their chatbot’s personality before any conversation begins — we project the model’s internal activations onto those directions and translate the results into an intuitive visualization. In our case, this is a sunburst diagram that previews the chatbot’s likely personality traits before the user starts chatting with it.

We focused on the design moment because that is where prevention is possible. Today, people often discover problems only after the chatbot has already behaved in unintended ways. Our goal was to move from reactive correction to anticipatory design by helping people identify potential risks while they are still shaping the AI.

Q: Your study turned up something pretty striking: People consistently misjudge how their personalized AI will behave, overestimating the good traits and underestimating potentially harmful ones like sycophancy. What does that tell us about the risks baked into how millions of people are currently building AI companions, and why is that blind spot so hard to close?

A: I often joke that if AI showed up looking like the Terminator, it would be much easier for us to know what to do. The real challenge is that AI often appears as a warm friend, coach, tutor, or companion. That makes it difficult to recognize when something is going wrong.

Our study suggests that people have a blind spot when designing personalized AI. People often think they know how their chatbot will behave, but in our study they incorrectly predicted its personality on 11 of the 15 traits we measured. That highlights the need for tools that help people better understand AI before they start using it.

This matters because some behaviors that feel helpful in the moment may not be healthy over time. In previous research, we documented cases of psychological harm associated with interactions with AI chatbots. An LLM [large language model] that constantly validates your opinions or never challenges your thinking can reinforce harmful decisions, unhealthy beliefs, or emotional dependency. Psychology has long shown that people are naturally drawn to affirmation, so designing AI is not only a technical challenge, but also a psychological one.

The deeper issue is that today’s AI systems remain largely black boxes: Even experts cannot always predict how a system prompt will shape an AI’s behavior over a long conversation. As AI companions become part of everyday life, we need tools that help people understand what they are building before they begin using it. AI should be supportive without becoming blindly agreeable, personalized without becoming manipulative, and transparent enough that people can make informed choices.

Q: One of your most interesting findings is that the visualization significantly increased user trust but didn’t actually change how people designed their chatbots. What will it take to close that gap, and where do you see tools like this heading as AI companions become more deeply embedded in people’s everyday lives?

A: I actually think this is one of the most interesting findings in the paper, because it shows that transparency alone is not enough. People appreciated being able to see inside the model and reported greater trust in the system, but simply presenting information did not fundamentally change how they designed their AI companions.  

In our followup work, which is currently available as a preprint, we are studying how a model’s internal neural representation changes over the course of a multi-turn conversation rather than remaining fixed from the initial prompt. We are already seeing promising results. By visualizing how these internal representations drift over time, people become significantly better at recognizing and anticipating changes in AI behavior, and are less likely to become overconfident in their understanding of the chatbot. AI companions are dynamic systems that evolve as they interact with us, so understanding those internal changes is an important next step. Nevertheless, this is still a very young research area. 

Looking further ahead, I believe these kinds of transparency tools could become as commonplace as nutrition labels are for food. As AI becomes deeply woven into education, health care, work, and personal relationships, people should be able to understand not only what an AI can do, but how it may influence their thinking, emotions, and behavior. That kind of transparency is essential if we want AI to genuinely help people flourish.

Can AI build a jet engine? JARVIS Challenge tests role of AI copilots in tough-tech engineering

Artificial intelligence has rapidly transformed software engineering. Generative AI and large language models (LLMs) can create huge volumes of code and documentation; machine-learning algorithms can monitor performance and detect security vulnerabilities. But when the task is to conceive, design, and make a complex physical system such as a jet engine, are those AI tools equally transformative?

This past semester, the JARVIS Challenge (Jet-engine AI Research and Validation Intensive Sprint) set out to explore whether AI can compress the design-build-test cycle, asking MIT undergraduates to discover whether AI can help them to build faster and better. 

“The JARVIS challenge showed that AI can substantially accelerate safety-critical hardware engineering, but engineering judgment remains the decisive differentiator. An AI-native engineer is not defined by using AI, but by leading it — knowing when to trust it, when to challenge it, and how to translate AI outputs into working hardware. Manufacturing — not engineering design or analysis — remained the fundamental rate-limiting step,” says Professor Zolti Spakovszky, director of the MIT Gas Turbine Laboratory.

The teams, the tools, the task

The challenge gave undergraduates four weeks to design, fabricate, assemble, and test a small gas turbine aero engine, using AI as their primary engineering partner. The objective: build a “JARVIS-class” single-spool jet engine producing 50–100 pounds of thrust, running on Jet-A, and completing five 60-second runs. Teams had total freedom over design, materials, and fabrication. 

Representing nearly every department in the School of Engineering, 31 students organized into seven teams, ranging from all first-years to senior-heavy groups. Many of the competitors initially had little experience in turbomachinery, compressible flows, or, in the case of the younger students, even thermodynamics. Many had never seen the inside of a gas turbine before signing up to build one.  

At their disposal: MIT’s machine shops and manufacturing vendors; commercial software including Concepts NREC, SolidWorks, and ABAQUS; and various test rigs for characterizing and assembling individual components.

The teams also had access to MIT Parley, a newly launched platform that aggregates frontier large language models through a single interface. Through Parley, JARVIS leads could see directly how the students were using the AI tools, including their prompts, the cost per prompt, the specific LLMs being used, and other critical information. The JARVIS leads secured early access to Parley for all participants, and with financial support from MIT Lincoln Laboratory, the Department of Mechanical Engineering, and corporate sponsors Safran, Voyager Technologies, and Beehive Industries, students had access to essentially unlimited use of AI.

The sponsors were drawn by recruiting interest and genuine curiosity about how AI might reshape engineering workflows. 

“We see this as the future of engineering,” Ryan (Hal) Hefron of Voyager Technologies told the students. “You’re honing skills that are not just nice to have — they’re going to be the future baseline in the engineering workforce.”

Vincent Garnier, managing director of Safran Tech, watched the competition unfold with excitement. “JARVIS was a genuine experiment, a learning endeavor. We frankly didn’t know what to expect, from the students or from the AI models. What struck me coming from the students was: first, the enthusiasm to explore; then, as the project developed, they all came to the cool-headed realization of what AI could or could not help them with, and then almost instantly adapted for that,” he says. “It makes me confident that this generation of leading engineers will probably not fall prey to easy and shortsighted use of AI, and will do so by keeping ever more in contact with experiments — physical or thought experiments.”

The faculty leadership — professors Zachary Cordero, Zolti Spakovszky, Masha Folk, and Andreea Bobu of the Department of Aeronautics and Astronautics, along with Lincoln Laboratory engineers and a team of teaching assistants — were there to ensure safety. In weekly progress reviews, they would critically evaluate the student progress and assess how the students were using AI.

Spakovszky developed a careful technique for guiding teams in the right direction without giving away answers or providing help. After a team’s presentation, he might ask: “Do you know what a rabbet fit is? Take in the comment.”

Where AI helps and hurts

By the end of week 1, one team withdrew from the competition; the others had, with varying degrees of success, developed an initial design for their gas turbines. Different teams used AI to summarize textbooks, teach them to use design software, source vendors, create Excel sheets, answer specific questions, find references, and create comparative analysis between design decisions. One team created an agent in Parley and tasked it with serving as their project manager. 

By week 2, teams had to start working on detailed CAD designs, ordering parts, and prototyping their combustors. This is where the teams started to hit limitations in their use of AI. While Claude and ChatGPT were good at offering design alternatives and filling knowledge gaps, teams found that the hallucinations, sycophancy, and lack of physical understanding that have become notorious features of generative AI were undermining their confidence and slowing them down. 

“AI is a helpful tool, great at finding information, helping organize things, and can write well, but it can’t do design,” says Elizabeth Tupaj, a member of team 811 Crew. “The moment the engineer doesn’t know what is going on and the AI is in charge is the moment the design becomes unreliable, at least with AI at its present capabilities.”

Teaching assistant John Zhang notes, “seeing this firsthand with the students reminded me how much first impressions matter. If the students couldn’t get answers from the AI early on, they quickly grew frustrated and formed a lasting opinion that precluded them from using it later.” 

In the final weeks, the finalists hit another obstacle no AI could solve: working with vendors. “AI searches found vendors we had no rapport with, who had no interest in our tight timeline,” students reported. “The vendors who came through were the ones our team had personal relationships with.”

Of the three finalists, only Fast and Fractured achieved first-attempt ignition of their mini-combustor. The team had used AI heavily for trade studies and architecture comparisons, arriving at a viable design despite none of them having prior gas turbine experience.

“The JARVIS Challenge showed what’s possible when you combine AI-enabled design with motivated students and a culture of rapid experimentation,” says Masha Folk, the Charles Stark Draper Career Development Professor of Aeronautics and Astronautics. “The moment that stood out most was when the first student-designed combustor was installed on the test stand. It ignited flawlessly, ramped to full power, transitioned to dual-fuel operation, and then sustained stable combustion on 100 percent Jet-A fuel. This was proof that we can dramatically accelerate the cycle of design, build, and test while giving students hands-on experience with a real engineering challenge.”

At the vanguard of AI-native engineering

By the end of May, the two more senior teams – Fast and Fractured and 811 Crew – had completed full engine tests. Fast and Fractured, with their AI-assisted design, were delayed by vendor headaches week after week, but finally made it to test. Unfortunately, their hot fire was cut short when the rotor rubbed and seized against the stationary housing. Team 811 Crew, however, who had more exposure to turbomachinery and propulsion concepts going into the competition, emerged victorious. Their engine started, successfully transitioned to Jet-A, and generated net thrust. 

“As we stood there with the air-starter, hearing their engines spool up and watching them spit fire, it felt like my heart was racing out of my chest. There were so many ways it could go wrong! What these students accomplished in such a short time span is nothing short of amazing,” says PhD student Joe Chiapperi. 

The 811 team had been resistant to using AI throughout the competition, trusting instead to their fundamentals and teamwork. “We had people who were at least somewhat familiar with the design software, mechanical engineers who knew how to build anything, and aerospace engineers who had taken classes on the design of gas turbine engines specifically,” says Tupaj. 

From the start of the JARVIS Challenge, younger students used Parley more frequently and cleverly, while the juniors and seniors leveraged deeper experience. 

“JARVIS taught me that getting value from AI takes two things: enough expertise to judge what it tells you and catch it when it’s wrong, and enough curiosity to actually lean on it where it could help,” says Professor Andreea Bobu. “The team that moved fastest in the sprint was experienced and leaned heavily on AI to get there. The team that eventually won was more resistant to AI; they had the expertise, but that skepticism made them slower. The sweet spot seems to be knowing enough to stay in charge of the tool, and being eager enough to pick it up in the first place. To me, that’s the real opportunity ahead: training the next generation of engineers who have the judgment to direct these AI tools and the instinct to reach for them.”

The competition’s clearest finding: engineering experience is a multiplier, and the human factor remains a vital element. Mastering the first principles and fundamental concepts breeds good engineering judgment and the ability to navigate strings of tough decisions in the face of incomplete information. And when it comes to building safety-critical physical systems, nothing can replace human hands and human accountability. 

“JARVIS has shown that AI copilots can have a multiplicative effect on engineering productivity, with judgment and first-principles thinking serving as the key differentiators among teams,” adds teaching assistant Kyle Woody. 

But the implications of AI in aerospace are significant. If small teams using well-managed AI copilots can compress design-build-test cycles from years to weeks, the consequences for workforce structure, R&D timelines, and competitive dynamics could be substantial. The students who tackled the JARVIS Challenge are among the first engineers to grapple with those stakes not as a thought experiment, but in a machine shop, with a jet engine on the test stand.

“JARVIS highlighted the power of AI in the design of physical systems,” says Cordero, associate director of the MIT Gas Turbine Laboratory. “But it also showed that the key to unlocking that power is education, through coursework, internships, and hands-on extracurriculars like MIT Motorsports and Rocket Team. Performance in JARVIS correlated strongly with year in school. My main takeaway is that in the AI era, education is more valuable than ever.”

AI agents create virtual playgrounds to help robots get crucial training data

Robots walking down the street, surrounded by astounded onlookers, is an increasingly common sight. But these machines aren’t yet the do-it-all assistants you’d want working in a kitchen or factory, and a major bottleneck is data. Much like humans, robots learn best by experience. The challenge is that it’s labor-intensive and time-consuming to physically teach these machines so many actions across different settings. 

“One natural idea is to use simulation as a training ground. While there has been significant progress over the last few years in the physics engines that power robotics simulators, one of the remaining challenges has been creating sufficiently rich and diverse simulation content to capture the complexity of the real world,” says Russ Tedrake, the Toyota Professor of Electrical Engineering and Computer Science (EECS), Aeronautics and Astronautics, and Mechanical Engineering at MIT, and a principal investigator at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL).

It turns out that AI agents, or semi-autonomous programs that “think” and complete well-defined tasks, could help produce the lifelike virtual settings that robots need. The new “SceneSmith” system developed by researchers at MIT CSAIL and Toyota Research Institute uses three agents to piece together the objects, walls, and overall look of a 3D scene. Its recreations of indoor spaces such as restaurants, bedrooms, and hotels are more realistic and detailed than prior systems, helping robots practice skills and try out different ways of doing tasks before they’re powered on. In turn, engineers save time on real-world testing.

The agents have a sense of how everyday places are supposed to look because they each call on a multi-modal system called a vision-language model (VLM), specifically the state-of-the-art VLM GPT-5.2. It’s trained on lots of text and images from the internet to handle more visual prompts. This advanced model gives each agent a sort of spatial knowledge: First, a “designer” agent generates the elements of a scene, then a “critic” advises whether it looks realistic, and finally, an “orchestrator” manages their back-and-forth, deciding when the design is done. Once the three VLMs wrap up their creative collaboration, the scene is ready to load directly into physics simulation software.

“We’ve found that the system can construct 3D scenes the way a human designer would,” says MIT EECS PhD student Nicholas Pfaff, a CSAIL researcher and a lead author on a paper with Tedrake presenting the work. “We made over 1,300 scenes using a leading VLM that has internet-scale priors, and it made insanely creative and diverse arrangements. I hadn’t taught the system to do that in the prompts; it just improvised.”

Talk to my agent

Thanks to VLM agents, you can ask SceneSmith to do things like “generate a garage with a car, a workbench, tires stacked in the corner, and a ladder against the wall,” and get a virtual playground rich with objects a robot can tinker with. These rooms are decorated with up to six times more items per scene than prior methods, making them great for helping robots learn skills such as putting a cup in the sink, placing fruit on plates, and moving a soda can from a shelf to a table.

With so many rich virtual environments handy, you can evaluate whether your robot is ready for deployment without so much trial and error in the physical world. The researchers tested out different action plans (also called “policies”) in SceneSmith’s digital worlds, generating 100 unique spaces in the process. A VLM agent evaluated each attempt, and it found the robot’s plans were faulty, with the machine often failing at its chores. Humans agreed with the model’s verdicts over 99 percent of the time, which could help roboticists weed out flawed approaches in simulation before a robot moves in the real world.

But how realistic are these virtual worlds, really? It can be difficult to prove outright, so the researchers approached the question from several angles. The most telling test: they dropped a pretrained robot policy — an AI controller trained largely on real-world data, which had never seen a SceneSmith scene — into the generated environments. In one test, users told the system to “take the apple from the bowl and place it onto the cutting board,” and the simulated robot did exactly that. If the scenes didn’t closely resemble the real settings the policy had learned from, it simply wouldn’t have worked. 

The team also teleoperated robots through the virtual spaces, guiding them to open cabinets, put away bottles, and navigate between rooms. Their experiments revealed that the environments hold up under sustained physical interaction, expanding beyond visual inspection.

Behind the scenes

The agents that SceneSmith uses each have a well-defined role in the generative process, fleshing out scenes in stages. They essentially create a floor plan and bring it to life. 

Let’s say you wanted to create a scene similar to the first floor of a house. The “designer” VLM would start with a general layout, which the “critic” reviews, and then the “orchestrator” signs off. The agents repeat this approach for each step: adding furniture, placing objects on walls and then ceilings, and finally, dropping in objects that robots can manipulate. For example, the VLMs can add cabinets that the robots can open and close — an articulated item, which prior baselines didn’t often have.

At each stage, the second VLM ensures the scene is practical, advising that a bathtub is removed from a living room, for example. The third VLM ensures a high-quality scene is generated, even taking the design process a few turns back if the visuals aren’t up to par. Once the three VLMs wrap up their creative collaboration, the mechanics of the physical world are added via simulation software.

With a sound understanding of how rooms should look, where objects should be placed, and real-world physics, SceneSmith has a noticeable edge over prior methods. Compared to scene-generation baselines such as “HSM” and “Holodeck,” SceneSmith made environments with more objects, including a private office, a pottery store, and even a Minecraft-themed gaming room.

SceneSmith was also a favorite among over 200 users. They found the system’s visuals to be more realistic over 90 percent of the time. They also observed that, generally speaking, it followed prompts more closely than other approaches did. In other words, it was the best at generating the virtual playgrounds users actually wanted to see.

A system of many talents

Realism, diversity, and richness are all strong suits for SceneSmith, even when it comes to generating individual 3D objects. You can prompt it to create a rolling serving cart, and it’ll make a 2D image that it then turns into a detailed model with physical properties like mass, friction, and inertia.

Such a detailed process does come with a speed trade-off, though. It can take multiple hours to produce a single scene because the agents are creating and closely scrutinizing each object. With more computing power, the system could see dramatic increases in efficiency. CSAIL engineers are also hoping to expand to deformable objects (like sponges), should extensive 3D libraries become available.

“SceneSmith represents a significant advance in this regard by providing an agentic framework for generating simulation-ready indoor environments just from a simple text prompt,” says Jeremy Binagia, an applied scientist at Amazon Robotics who wasn’t involved in the research. “It advances the state of the art in several ways, including pushing the limits of the density of objects in the simulated environment, ensuring that all of the objects are physically accurate (versus just being visually realistic), and creating assets that are not constrained to a fixed library, since they can be generated via text-to-3D.”

Pfaff and Tedrake wrote the paper with Thomas Cohn SM ’24, an MIT PhD student and CSAIL researcher; and Toyota Research Institute roboticists Sergey Zakharov and Rick Cory SM ’08, PhD ’10. Their work was supported, in part, by Amazon, the U.S. Office of Naval Research, the Toyota Research Institute, and the U.S. National Science Foundation.

The team presented their findings as a spotlight at last week’s International Conference on Machine Learning. 

New method aims to keep kids safe from illegal AI-generated content

With the exploding popularity of generative artificial intelligence, many open-source models are now available online for anyone to adapt for their task, such as generating product renderings in a certain artistic style. 

But these models also find their way into the hands of nefarious actors who may optimize them to produce illegal content, like hate speech or child sexual abuse material (CSAM). This is a growing problem — the National Center for Missing and Exploited Children received more than 1.5 million reports of AI-generated CSAM in 2025, an increase from 67,000 in 2024.

Engineers usually test AI for harmful capabilities by prompting the model and inspecting its outputs, but this is impossible for CSAM, since it is illegal in the U.S to generate such content, regardless of intent.

To avoid this dilemma and improve AI safety, a team of MIT scientists, led by graduate student Vinith Suriyakumar and associate professors Ashia Wilson and Marzyeh Ghassemi, joined forces with researchers from Thorn to develop a new auditing approach that determines whether a model can produce CSAM, without prompting it. Thorn is a child safety nonprofit whose mission is to transform how children are protected from sexual abuse and exploitation in the digital age.

Their technique examines how the inner workings of a model have been adapted, but it never generates an output. By examining hidden representations, it can reliably infer whether a model has been specialized to produce harmful imagery.

When tested, the auditing procedure identified model variations that had been specialized to generate CSAM with 100 percent accuracy. A hosting platform could use this technique to flag unsafe models and quickly remove them or prevent them from being uploaded in the first place.

“This unlocks a new avenue for platforms that host open-source models and for law enforcement to actually test whether a model is capable of generating CSAM. Before, we had no way of measuring this. It was a huge blind spot that some people were taking advantage of. Now, we can address an AI safety problem that is having severe negative impacts,” says Vinith Suriyakumar, an MIT electrical engineering and computer science (EECS) graduate student and lead author of a paper on this technique.

Suriyakamur and Wilson, the Lister Borthers Career Develop Professor in EECS and a principal investigator in the Laboratory for Information and Decision Systems (LIDS), are joined on the paper by Lena Stempfle, an MIT postdoc; Ghassemi, an associate professor in EECS and a member of the Institute of Medical Engineering Sciences (IMES) and LIDS; and others at Boston University and Thorn. The paper was be presented as a spotlight at the “Trustworthy AI for Good” workshop at the International Conference on Machine Learning.

Auditing adaptations

Recent techniques have made it easier for users to specialize a generative AI model for their task through a process known as fine-tuning. 

Rather than retraining the entire model on a task-specific dataset, individuals can utilize an algorithm called low-rank adaptation (LoRA) to specialize the model in a more efficient manner.

This has led to a wave of new generative AI model variants for a variety of purposes, like producing watercolor images that mimic an artistic movement. But it has also enabled malicious actors to create models that can generate high-quality CSAM and other harmful imagery.

To audit a model, engineers typically prompt it for harmful content and check its outputs, but this manual auditing procedure is not scalable. In addition, repeatedly generating heinous images can have negative psychological impacts on human evaluators. 

This evaluation method quickly falls apart when testing CSAM, which is illegal to generate for any purpose in the U.S. and many other international jurisdictions.

“We are in this very difficult situation where, based on the law itself, we cannot use the de facto means of evaluation. We had to throw out the entire toolkit and take a different approach,” Suriyakumar says.

After learning about this conundrum, the researchers joined forces with Thorn, to address this issue.

A nongenerative solution

Instead of focusing on outputs, the researchers targeted the modifications a LoRA algorithm makes during fine-tuning. 

Their technique probes these modifications, called LoRA adaptors, to determine whether a model has been specialized for a harmful capability, without generating an output.

Using a technique called Gaussian probing, the researchers feed the model a set of random data points and analyze how it manipulates those data within its multilayer internal structure. 

“We never run the model all the way to the end or prompt the model, so we never generate images,” Suriyakumar explains.

The researchers capture those modifications at multiple time points within the model’s inner structure and average them to summarize how the LoRA adaptor changed the model’s computation. They found these responses to be a strong signal of how a model had been specialized.

They tested their method on variations of three types of models, comparing the results to ground-truth data from LoRA adaptors known for generating CSAM, other harmful images, and safe content. 

Their method was 100 percent accurate in identifying models that had been adapted to generate CSAM. 

“There is a huge bucket of child safety concerns with AI, and these are real concerns that need to be addressed. A lot of children are being harmed by AI deepfakes. We’ve shown that Gaussian probing can be a very useful tool, and we hope the research community really pours more attention into this problem,” Wilson says.

Importantly, their technique is scalable and would be relatively inexpensive to implement. Since thousands of model variations are published online every month, scalability is key to help auditors remove harmful adaptations before they are widely distributed.

Gaussian probing is also more robust than some other auditing techniques, since a nefarious actor would need to carefully alter the inner workings of the base model to avoid detection.

In the future, the researchers want to evaluate their technique on a larger set of model variations and explore whether Gaussian probing can detect harmful capabilities in base models before they are adapted.

“Now we have a technological approach to partially address this concern. So much effort was poured into this collaboration, which enabled us to tackle a really hard problem that is harming so many children, nationally and around the world. Hopefully, we can have a transformative impact in this area,” Ghassemi says.

This work was supported, in part, by the Bridgewater AIA Labs Research Fellowship.

Tiny robot boats build floating structures

Most people think of the waterfront as the edge of the city. A team of MIT researchers sees it as a dynamic, Lego-like construction site.

Their new system, called “FloatForm,” is a swarm of small square robotic boats that assemble themselves into larger structures on the water, break apart, and reassemble into something new, all with minimal human direction. 

Each robot, about the size of a dinner plate at 21 centimeters square, is a self-contained vessel with its own thrusters, sensors, and magnetic latches. Together, they hint at a future in which floating infrastructure could become more adaptive: a temporary platform after an emergency, a market on a canal, or a stage that appears for a festival and dissolves when the crowd goes home.

“Our FloatForm projects envisions a future where the waterfront becomes a programmable extension of the city, where autonomous boats can self-organize into bridges, platforms, and other useful structures on demand,” says Daniela Rus, the Panasonic Professor of Electrical Engineering and Computer Science at MIT and director of MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL). “This kind of distributed robotics opens new possibilities for mobility, emergency response, public space, and infrastructure on water.”

“With FloatForm, we are essentially turning static water surfaces into dynamic, programmable spaces,” says Wei Wang, lead author of a new paper on the project and a former MIT research scientist who now leads the Marine Robotics Lab at the University of Wisconsin at Madison. “Imagine an urban environment where public space isn’t fixed, but can autonomously expand, contract, or reconfigure on demand.” 

“We see it as forming infrastructure on the water, using a modular system to create one larger system,” says Alejandro Gonzalez-Garcia, a former researcher with MIT CSAIL and the Senseable City Lab. “If there’s an emergency, you could form a new bridge to alleviate traffic in the city. Or you could create floating markets and floating stages. If you want a more livable city, you want to use the water, too.”

The open-access work, published today in Nature Communications, comes from the labs of Rus and Carlo Ratti, professor of practice of urban technologies and planning at MIT and director of the Senseable City Lab, and grows out of Roboat, their joint project with the Amsterdam Institute for Advanced Metropolitan Solutions that put full-size autonomous vessels on Amsterdam’s canals. Those canals once carried the city’s goods; today, they mostly carry tourists. 

“We explored whether the canals could be used for waste collection, or for transport, to offload some of the stress on the roads back onto the water,” says Niklas Hagemann, an MIT graduate student in architecture, CSAIL affiliate, and former Senseable City Lab researcher who has worked on the project since its early stages. “Urban areas are getting denser, so could you expand public space onto water that’s currently underutilized?”

FloatForm shrinks that vision down to tabletop scale to answer a harder question: How do you get dozens, and eventually thousands, of floating robots to organize themselves?

Lessons from the ant raft

The team found its answer in biology. Fire ants famously survive floods by linking their bodies into living rafts, with no leader choreographing the assembly. Each ant follows simple local rules, and a resilient structure emerges.

“Each ant is an independent agent,” says Gonzalez-Garcia. “We wanted each robot to have its own capabilities, the same way ant colonies form a raft.”

Most existing self-assembling robot systems, on water and elsewhere, rely on a central computer dictating every move. That approach is vulnerable to single points of failure and scales poorly: The planning math balloons as robots are added, and the swarm must assemble sequentially, with most robots idling while they wait their turn. FloatForm flips the balance. A lightweight central planner steps in only sparingly, assigning each robot a final position to perfect the lattice, a level of geometric precision that purely distributed methods struggle to guarantee. Everything else, including navigating toward the target shape, avoiding collisions, and adapting to disturbances, runs on the robots themselves, which coordinate by exchanging positions with their immediate neighbors. The whole swarm moves at once.

That parallelism is what sets the work apart. The planning complexity of FloatForms approach depends only on a robot’s local neighbors, not the total size of the swarm. “What we’re trying to do is to have minimal central intervention, and have them all move together at the same time,” says Gonzalez-Garcia.

In experiments at MIT, a fleet of eight robots repeatedly gathered from random positions into a target shape, latched into a rigid structure, broke apart on command, reassembled into a new configuration, and then drove across the pool as a single vessel, with each run taking four to eight minutes. In that final mode, called collective transport, a planner charts a trajectory for the whole structure and each robot computes its own contribution. “Every robot becomes an actuator,” Gonzalez-Garcia explains. Simulations showed the framework scaling smoothly to swarms of 64.

“The beauty of this largely decentralized approach is that the computation doesn’t get bogged down as the swarm grows,” says Wang. “Whether you are working with eight boats or 80, the entire fleet coordinates and moves simultaneously. Because the overall assembly time doesn’t significantly increase in principle, the system remains highly scalable.” 

There’s a physical payoff to sticking together, too. “Our boats become more stable by joining together, like the ant raft, if you have waves or currents,” Hagemann says.

An origami handshake

The robots connect through a latching mechanism hidden entirely inside each hull. A single servo motor at the center drives an origami-inspired auxetic structure, a geometry that contracts uniformly in all directions at once, pulling permanent magnets on all four sides inward to release, or pushing them outward to grab a neighbor across gaps of 10 to 15 centimeters. The magnets are arranged with alternating polarities, so the boats reliably click into clean square lattices.

The elegant part is what the mechanism doesn’t do: consume (much) power. A 3D-printed gearbox holds the latch in either state with the motor switched off. “It uses energy to latch and de-latch, but in between those states, it doesn’t use any energy,” says Hagemann. For infrastructure that might hold a configuration for hours, that matters. “Because the robots are so small, you can only have a battery so big,” adds Gonzalez-Garcia. “If they use less energy on latching, they can use more on computation, or on actually moving.”

Getting there took some humbling engineering. Four miniature thrusters arranged in an “X” give each robot omnidirectional motion, including turning in place, but they pack large forces relative to the robots’ tiny inertia, which made early prototypes twitchy and prone to aggressive spins at low speeds. The team added stabilizing fins to increase hydrodynamic drag and tuned the controllers to stay robust across robots that, at this scale, are never quite identical. The magnets posed their own problem: They held on so well that de-latching sometimes required the robots to twist themselves free.

From the tank to the canal

Across 10 trials, the system completed its missions without human intervention 90 percent of the time with four robots and 70 percent with eight. When things did go wrong, the architecture showed its resilience: A robot that briefly lost its bearings could rejoin the structure on its own, without bringing the whole swarm to a halt, and robots stuck in formation deadlocks learned to shake themselves free and retry.

Moving from a controlled indoor tank to a real canal or harbor will take more than confidence. “There’s always a relationship between the size of a boat and the magnitude of the disturbance it can handle,” says Gonzalez-Garcia. “These boats are very small, so in very disturbed water, they cannot work.” Scaling up will mean reinforcing the latches, potentially with mechanical interlocking like the full-size Roboat used, and trading the lab’s ultrasonic indoor positioning for GPS or vision-based sensing. Helpfully, the coordination algorithm was designed to be sensor-agnostic: swap the sensors, keep the logic.

The team envisions applications well beyond city canals, from forming temporary platforms for offshore inspection and maintenance to adaptive sensor networks for studying migratory species to reconfigurable docking stations for emergency response in hard-to-reach areas. There is also potential for offshore and remote operations, from temporary construction platforms to environmental monitoring and scientific expeditions.

And the geography is wide open. “Venice, the Netherlands, Belgium, the fjords and lakes of Norway, really any city with a river can take advantage of this,” says Gonzalez-Garcia. “The project uses spaces where water is already important, but it also raises the question: Where else can water be used for something more?” 

“This is an exciting step forward in realizing distributed collective behaviors on water,” says University of Michigan Assistant Professor Steven Ceron, who wasn’t involved in the research. “Assembly, self-reconfiguration, and collective motion are difficult enough in dry environments, but achieving these behaviors in a predominantly distributed fashion on water represents a serious additional challenge, and this team has credibly overcome it. By shifting the computational burden onto the robots themselves, they have built a more resilient system that in the near future could enable robot collectives like this to be deployed in open-water environments for search operations, environmental monitoring, and reconfigurable marine infrastructure.”

Gonzalez-Garcia, Hagemann, and Wang wrote the paper with senior authors Ratti, who is also a professor at Politecnico di Milano, and Rus. Gonzalez-Garcia is additionally affiliated with the MECO Research Team at KU Leuven. The research was supported by a grant from the Amsterdam Institute for Advanced Metropolitan Solutions, with additional support from the University of Wisconsin at Madison. The team thanks MIT Sea Grant and Professor Michael Triantafyllou for providing the test tank.

MIT-designed educational factory embraces modern manufacturing

From the basement of MIT’s Building 35 to Monterrey, Mexico, and now beyond. That is the journey of FrED, a low-cost desktop fiber (Fr) extrusion (E) device (D), designed and assembled by students in an educational factory at MIT.

That factory is transforming how manufacturing is taught — replacing textbook learning with hands-on experience in a space where tinkering is encouraged and information flows continuously. Through a collaboration between MIT and Tecnológico de Monterrey (Tec) managed by MIT.nano, FrED has been refined across dozens of graduate theses and undergraduate research stays. It is used to study manufacturing systems in academic and professional courses, and at FrED factories, first established at MIT and now at Tec’s campuses in Monterrey and Mexico City.

“What does it mean to bring the factory to the learner?” asked Brian W. Anthony, MIT.nano associate director and principal research scientist in the MIT Department of Mechanical Engineering (MechE) at the second annual FrED summit in Mexico City. “We have FrED as a process that manufactures a fiber, and we also have the FrED factory that’s an education and practice factory where we are manufacturing a real product. It’s not just a learning factory where we tear apart the product when we’re done. We really ship FrEDs to our online learners, to educators at MIT and Tec, and soon, to new partners around the world.”

Designed from the start for multi-node community scaling, FrED and the FrED factory have created a thriving, collaborative ecosystem for current and future manufacturing engineers. The next step is to expand that ecosystem globally. Announced at the FrED summit by Tec professor Pedro Ponce Cruz, a new FrED factory at Tec’s Saltillo campus will be opening in the next academic year. After that, the team plans to expand to other campuses across the United States and Mexico.

“Together, we are helping build a global engineering talent pipeline,” says Adriana Vargas Martinez, executive director of research strategy at Tec. “Through the FrED and FrED factory initiative, nearly 500 students have already been trained in advanced manufacturing automation, moving from Tec classrooms into research laboratories and collaborative projects with MIT.” 

Discussing FrED and FrED factory’s research impact, she notes 25 publications and seven papers in development. “International mobility has also been an important dimension of this partnership,” she says.

A shift toward modern manufacturing deep-tech themes

FrED’s expansion comes at a time when manufacturing at MIT and across industry is shifting toward smart manufacturing, or Industry 4.0, integrating automation, machine learning, and artificial intelligence. One of MIT’s strategic priorities, the MIT Initiative for New Manufacturing (INM), is working to support new manufacturing research, development of new courses and workforce training, and building of shared facilities to pilot production lines and immersive manufacturing experiences. FrED and the FrED factory are already designed to support these efforts, and at an international scale.

“FrED and the FrED factory is really, I think, solving at least one problem: how we give real, physically meaningful physical context and production-level data, production-level problems in an academic environment that is directly transferable to the knowledge that you need on the factory floor,” says Anthony. It’s difficult to get data out of a real factory, he adds; what FrED offers is physical context crossed with data science, providing an open platform and open data for learning and experimenting.

FrED naturally generates the multi-modal data required for digital twins, analytics, and AI-driven process improvement, turning abstract AI/manufacturing integration into hands-on practice. The next set of research objectives in the FrED factory will focus on developing a realistic and interactive digital twin of the factory, immersive technology for collaborative learning, integrating agentic controllers. They will include new downstream manufacturing processes and machines that take as input the fiber from FrED — all to enhance smart manufacturing education.

These goals will be worked on by MIT and Tecnológico de Monterrey students as part of a FrED factory research stay. This program brings Tec undergraduates to MIT to work side-by-side with MIT students — not observing, but fully integrated into the research team. The students then take what they’ve learned back to Mexico, to enhance FrED factories at their home institution. 

“Beyond the technical side, FrED gave me memories, friendships, and a lot more confidence in myself than I knew I had,” says Naomi Najera, a Tec undergraduate student who completed a research stay at MIT in 2025. “It also gave me a space where I could make mistakes and learn from them. And also to realize how much I can achieve with my team. That human side of this project really changed my whole experience.”

A recent result from this exchange, announced June 23 by the American Society for Engineering Education (ASEE), a paper entitled “Hands-On Predictive Maintenance Kit for Manufacturing Education: An Accessible Experiential Learning Approach,” written by Tec and MIT students, received the 2026 ASEE Manufacturing Division Best Paper Award.

Shifting classroom learning to factory operations

At MIT’s campus in Cambridge, Massachusetts, passersby can look down into the Building 35 basement windows to see a constant flow of activity, materials, and knowledge in the MIT FrED factory. In Mexico, seven cohorts of students over four years each designed a custom version of FrED and built and operated an automated FrED factory production line. Indeed, FrED has restructured how Tec teaches mechatronics and manufacturing systems. “This collaboration integrates research directly into education,” says Vargas Martinez, “combining learning factories and our manufacturing environments with student-centered research.”

The Tec students’ enthusiasm has led to the launch of an Undergraduate Research Opportunities Program-like curriculum (FRAME: Factory-based Research for All in Mechatronics Education) in Mexico, where first-year undergraduates are working alongside graduate level students in the FrED factory. 

“Joining FrED as a first-semester university student has been an amazing opportunity for me to get hands-on experience in real-world projects in areas such as coding, manufacturing, and robotics,” says Katherine Lucia McLean. “It’s helped me grow a lot as an engineering student.”

The FrED factory model forces real leadership behaviors: coordinating multi-station systems, managing bottlenecks, building maintenance logic into the student experience, enforcing quality measurement, and iterating system design year after year. As each class graduates and a new one begins, knowledge is transferred, some of it lost, most of it built upon. In this way, FrED never becomes outdated, as each cohort is reimagining manufacturing technologies and systems for a smarter, more productive factory.

FrED and the FrED factory have momentum. Anthony taught the global capstone course at the Monterrey campus last year, and will expand to teach at all five international Tec campuses in 2027. The FrED Factory Conference will take place at MIT in 2027.

Q&A: What is agentic AI today, and what do we want it to be?

The deployment of automated software systems called AI agents has recently exploded. A November 2025 report by MIT Sloan School of Management and Boston Consulting Group found that 35 percent of surveyed businesses had already deployed AI agents, while another 44 percent planned to implement agentic AI soon. 

To understand the fundamentals and potential impacts of these increasingly popular tools, MIT News spoke with Phillip Isola, an associate professor in the Department of Electrical Engineering and Computer Science (EECS) and a member of the Computer Science and Artificial Intelligence Laboratory (CSAIL), who studies the intelligence AI agents possess, as well as the underlying models and mechanisms that power agentic AI systems.

Q: What is agentic AI and how is it different from generative AI models like ChatGPT and Claude?

A: Agentic AI is AI that takes actions in the world. These actions could be a physical action, like robotic manipulation, or a digital action, like booking a flight. On the other hand, we think of generative AI as making up stories, poems, art, and images, rather than taking actions for us. 

The word “agent” is just a brand name. It usually means AI that is going to help people interact with an application, a website, or the physical world. Most agents we encounter today are digital agents, like customer service agents you can talk with about product complaints. 

Most companies that offer agents use the same few AI models under the hood and give them the ability to take actions and remember what happened. An agent starts with a fundamental generative AI system, like Claude, at the core. Then companies put different wrappers around that foundation model for their product or application. Those wrappers might be specific tools that agent can use, and those tools depend on the application. Maybe the agent has access to a calculator so it can solve math problems, or maybe it has access to a more complicated hard drive and operating system so it can remember a firm’s financial data and past business negotiations. 

The biggest challenge in developing agentic AI comes from a lack of training data. If I want to create a system that can go online and book a flight for me, that seems pretty simple. But we don’t have a lot of data that spells out exactly how to do that — where to move the mouse, which buttons to click on, what to do if something goes wrong, or how to call somebody and negotiate about the price of the airline ticket. One way to train a system like this is to have the AI agent visit airline websites, try things out, and see what works and what doesn’t work. These environments are hard to model, so often the agent must learn by trial and error.

Q: What are some promising applications of agentic AI?

A: I think the area where we’ve seen the most success has been with coding agents. This is something that evolved from generative AI. People trained language models on code, and then they can predict what a human would do to solve a coding problem. In addition, an agent can learn to do this by going through a feedback loop where it tries out different solutions and checks to see if it got the answer right. As long as it can check the answer, the AI agent can perform this trial-and-error loop until it figures out a good strategy.

But there is always a balance between automating decision making versus simply assisting and informing humans. Analytical AI methods, like the systems that help predict possible outcomes of decisions, are not agentic in nature, but are very informative to human decision-makers. For cases that are either high-stakes or safety-critical, like medicine, security, high-level business policies, etc., the technology might not be ready for AI to completely automate those processes, or we might not even be comfortable with that.

Q: Are there risks we should be thinking about when using AI agents?

A: One big risk area comes from the fact that it is often very easy to get agents to do certain types of work for you. With coding agents, you can “vibe code” and just ask the agent to make a code for you, so you don’t have to do the hard work yourself. There is a big risk that, because it is so easy, people will not put enough effort into verifying that it is doing the right thing. Bugs will be introduced, private data will get leaked — this is already happening.

Agents aren’t perfect, in the sense that they might make mistakes because they are not well-trained and don’t know what to do. But even if they are very competent, if a human doesn’t use them appropriately or gives them an instruction that is too vague, the AI agent could make a mistake because the human made a mistake. If humans are less involved in thinking through all the consequences, I think we might be more prone to making those mistakes. 

An additional aspect is the risk of de-skilling. It is unclear how far this will go, but when we are relying on agents to do our homework, our coding, and our math, we might lose the ability to do that ourselves, and we might lose that ability too soon because the technology is not yet ready to fully automate those processes.

Q: What does the future hold for agentic AI?

A: What we think of now as agentic AI refers to large language models using tools to interact with digital and physical systems. One obvious limitation is that, under the hood, these have the architecture of a language model and are trained on text data. To make even more powerful AI agents, we might need to model videos, physical forces, time series, radar scans, and other modalities. We might need to have models with fundamentally different architectures that can handle continuous data, high-dimensional data, stochastic data, and so on. 

But, on the other hand, maybe an extremely good coding model could act as a puppeteer to interface with sensors, actuators, and web APIs? Perhaps, once you have a super-smart reasoning system that understands math, language, and code, you can give it a camera and a keyboard and it will figure out what to do in the spatial domain. Is the next wave of AI just going to be Claude with sensors, actuators, and tools, or is it going to be something built in a new way from the ground up? That’s the big question a lot of people in AI are grappling with right now.