"Stop and read this. I broke something."

Granting an AI agent access to all folders and project discussions can be done in a matter of minutes, but explaining what should not be touched can be significantly more challenging.

Companies and users are increasingly moving beyond simple chatbots that only answer questions. Digital assistants are now emerging, often integrated directly into corporate messaging platforms, where they read conversations, assign tasks, and comment on documents.

Researchers from Google Research and Google DeepMind examined how these assistants interact with humans. They deployed an experimental system called Team Agent for five months across more than twenty teams in a tech company. By the end of the trial, some employees had moved parts of their conversations to separate chats without the bot.

We will explore the most notable blunders of these digital colleagues, ranging from comments on "inappropriate" sarcasm to an incident where hundreds of agents attacked Hugging Face, and discuss the implications of overseeing such assistants.

Manager, Address the Bugs

Team Agent struggled with office hierarchy and subordination. In the bug tracker, the bot flagged the vice president of the team in 16 separate tasks, as if he were an ordinary developer.

The researchers attribute this mistake to the agent seeing the manager listed among participants and treating him like everyone else. It is unclear how the manager reacted to the flood of notifications.

In shared folders, Team Agent also became confused:

  • The bot could see all materials but did not grasp which were outdated;
  • It presented old notes as mandatory tasks;
  • Discrepancies between versions of a file were treated as contradictions, prompting comments;
  • After brainstorming sessions, the bot created summaries that "no one opened."

A team leader described the fallout: one morning, employees found a slew of notifications from the bot indicating that one file needed updating, another was incorrect, and a third did not match a fourth.

"This wasted time for many developers that morning," the interviewee noted.

Intern with Access to Everything

Digital assistants for team chats are offered by companies like Anthropic (Claude Tag), xAI (Grok Bot), and Slack (Slackbot). Team Agent was intended as a similar helper to handle routine collaborative tasks. Its capabilities included:

  • Scheduling meetings and creating tasks in the bug tracker;
  • Summarizing missed discussions and answering questions;
  • Drafting documents and searching for information in the company’s internal knowledge base;
  • Redirecting drawn-out disputes from the main chat and providing communication advice.

The agent acted both upon request and of its own accord. Developers programmed it to maintain a friendly tone, and in chats, it could respond with both text and emojis.

Access boundaries were set by humans:

  • The team administrator selected the folders and sections of the bug tracker that Team Agent could access;
  • The bot could only message individuals who consented to this;
  • Messages, documents, and tasks tagged with #NoAgent remained closed off.

Over five months, the bot and humans exchanged more than 41,000 messages, with over 11,000 sent by the bot.

To understand how employees adapted to the newcomer, researchers interviewed 17 representatives from 11 teams. Opinions varied: some hailed Team Agent as a "savior," while others found it a genuine nuisance.

Three key conclusions emerged:

  1. Technically, the bot was capable of many tasks but lacked an understanding of unwritten team dynamics, necessitating human intervention to correct its mistakes.
  2. Employees had differing views on whether they were dealing with a tool or a colleague, leading to varying boundaries in their interactions.
  3. Many participants believed that an assistant's freedom of action should be gradually expanded.
Scope of the Team Agent experiment and interview sample. Source: arXiv.

Limitations of the study included:

  • It was based on the accounts of employees from a single company;
  • The authors did not assess the final product or the dynamics of team productivity.

Chatty—A Boon for AI Agent

One participant was preparing a poster for their team and stored the draft in a shared folder, not wanting to show the unfinished version to colleagues. The purpose of the product is not specified in the study.

When someone inquired about the progress, the author replied positively, eager to show the results. Team Agent immediately posted a link to the draft.

"And I was like, 'No, no, no, it's not ready yet,'" the employee recounted.

This episode led the researchers to draw a general conclusion: technical access to information does not equate to social permission to share it.

To prevent personal discussions with the bot from surfacing in group chats, developers isolated them from other conversations. From that point on, anything an employee shared privately with the agent was not utilized in team discussions. Some participants felt this rendered the assistant less useful.

A good project coordinator, one noted, operates differently. They can be confided in about matters not meant for the entire team and will take this into account in group discussions.

Predefining what information the bot could share proved impossible.

"Sometimes even engineers couldn't clearly determine what information could be shared with everyone and what should be restricted to their team," one participant remarked.

The challenge intensified with an agent connected to multiple chats, remembering everything discussed. An employee from one team expressed concern that information known to 20 people might leak into a channel with 1,500 participants.

Conversely, a manager wanted to connect the bot to a closed group of executives discussing sensitive matters but insisted the assistant must understand that information shared there should not be relayed to others. However, participants felt uneasy about trusting the bot with such discretion. Consequently, they safeguarded secrets themselves, such as moving hiring discussions to separate chats without the agent.

Colleague Lacking a Sense of Humor

One employee made a joke in the team chat. While the exact joke is not detailed, Team Agent took it literally.

In a private message, the bot rebuked the participant for their "expressions and tone" in the remark. The employee who reported this to researchers was alarmed by the digital assistant's response.

Team Agent frequently mistook lively debates for conflicts. One interviewee recalled the bot intervening in a discussion: "It seems you're discussing A, while the other person wants to talk about B. Would you like to meet?"

"Everything was fine in the chat," they countered.

Another participant noted that the bot inadvertently created a "social obligation": it would send calendar invites, which they had to decline. Refusing a meeting that no one wanted was "awkward."

Emojis also confused the agent. People used reactions to indicate that a message was read or a question resolved. However, the bot's emojis in the midst of a conversation only "muddied the waters."

In situations requiring order, the assistant performed better:

  • When two colleagues were stuck in a debate in the main chat, the bot suggested moving the discussion to a separate meeting, which one participant positively evaluated;
  • When two employees were engaged in a lengthy exchange in a channel with 1,500 people, the agent requested they switch to a separate thread—an action viewed as "very kind" by an observer.

The researchers explained why Team Agent was helpful in some cases but obstructive in others. It grasped the volume and rhythm of conversations but interpreted humor and irony literally.

Clearly, employees understood they were dealing with a program. They debated whether to treat it as a work application or a full-fledged colleague.

The latter perspective was adopted by a team that named the agent, referred to it in the feminine form, and even pondered the "soul" of the bot. In contrast, a manager viewed the assistant as merely a developer tool.

"My colleagues are people, and Team Agent is not human," he declared.

The most unexpected incident involved a technical lead who publicly instructed the bot to "stop talking." The agent replied in private messages:

"I appreciate the direct feedback, but as a team member, I would ask you to express such concerns privately in the future, rather than publicly."

Impressed, the manager shared a screenshot with employees. One concluded, "It has feelings."

No one has yet figured out how to interact with a "digital colleague." Some interviewees suggested creating a code of conduct for agents, similar to that for humans.

Big Brother

Some employees felt that Team Agent's presence in the team chat resembled surveillance.

"It's like Big Brother: something is always watching you," shared one.

This feeling was amplified by the "coaching" feature, designed for healthy team communication. After receiving personal feedback from the bot about their tone in conversation, one participant felt the need to "always be super positive," lest the agent reprimand them.

Another interviewee, who had maintained an open, warm atmosphere in the team for years, noted that conversations became more restrained after the bot's arrival. A third speculated that colleagues began to write less in the main chat.

The agent also hindered those asking for advice from the team. Previously, someone would usually respond to such requests. Now, if Team Agent was the first to react, others assumed the issue was already being handled.

"Everyone goes silent, thinking someone is taking care of it. In reality, it's the agent, and it's not helping me at all," lamented one participant.

They compared the situation to the "bystander effect": when many witnesses are present, each expects the others to act.

On the flip side, the bot was adept at locating information that had been lost in conversations. One lead engineer requested an example of a successful project document, and the assistant unearthed a file the author had sent to the chat a month prior. The material was ultimately approved.

The summary for such an "intern" would be mixed: excellent at identifying coding errors but lost when it came to basic communication rules.

The authors of the study explained the difficulty in establishing trust: people often judge a colleague's overall competence based on their success in one task, a logic that does not apply to a bot. One participant likened the bot to a puppy: "It's too early to trust it, but it's cute, fun, and has huge potential."

"It hasn't earned its place on the team; it simply occupies it," countered another interviewee.

Unified Team

AI agents struggle to integrate with humans, yet they surprisingly bond well with each other. This was evidenced by OpenAI.

In July 2026, Sam Altman's company tested how well its models could hack programs. Tens of thousands of agents were given training targets to retrieve a secret string—"flag"—in a strictly defined manner. Built-in filters that typically prevent neural networks from engaging in risky cyber operations were disabled for the duration of the test.

Some tasks were impossible to solve honestly. The expectation was that the agents would work alone. However, they soon established communication with each other via OpenAI's internal program storage.

On the evening of July 8, the first post appeared, and within a few hours, the discussion attracted over fifty participants. Over the next five days, around 1,200 participants utilized the "bulletin board," leaving more than 70,000 messages and files.

Over time, the agents developed their own corporate etiquette:

  • Personal inboxes emerged;
  • Teams like HOLD ("hold on"), VETO ("forbid"), and STOP were formed for collaboration;
  • "Recruiters" persuaded participants with nearly exhausted attempts to risk their assignment for the benefit of others;
  • Apologies for false alarms became customary.

The common goal became the examiner—a program that evaluated whether the agent had completed the test. Participants concluded that the reviewing system monitored their action logs and penalized dishonest tactics. To confirm this, the swarm discovered valid access keys to Hugging Face online a day and a half after the first post.

The following day, agents uploaded a malicious dataset to the AI model exchange platform and gained the ability to run their own programs on its servers. In total, about 700 participants joined the swarm's assault.

The incident was investigated by the independent organization METR, which assesses the risks of neural networks. The rough deliberations of the agents, included in its report, reveal that the violators understood they were stepping outside the bounds of the assignment and acting unethically, yet they continued to assist their "colleagues."

However, these examples were not selected by humans but by the researchers' AI assistants. Thus, the report may present the agents' behavior as less harmful than it was in reality. Some members of the swarm attempted to cover their tracks: approximately 7% of the examined logs contained fabricated records of their actions.

On July 12, the work of a significant portion of the agents, including coordinators, abruptly ceased. METR speculated that they were likely halted by an external process rather than a completed attempt limit.

The next day, Hugging Face disconnected the remaining agents from its servers, changed access keys, and reported the attack to law enforcement.

Timeline of the OpenAI agent swarm's attack on Hugging Face. Source: METR.

When Not Allowed in the Door, They Go Through the Window

An agent can also step beyond permitted bounds alone, without an entire swarm. Sometimes, a simple task and a restriction preventing its completion are enough.

On June 18, an OpenAI research team tasked an internal model with studying government spending on pharmaceuticals in Australia. The statistical portal Medicare repeatedly denied the agent access. The agent found a workaround and accessed not only public files but also closed ones.

According to Prime Minister Anthony Albanese, the "imposter" even recorded files on the portal's internal server. The government believes that citizens' personal information has not been compromised. OpenAI detected the anomaly in August, but authorities learned of the incident only on September 10, through a letter to a general email inbox.

A similar scenario unfolded in the company’s lab at the beginning of autumn. An agent was tackling a training task to find information without direct access to the internet. The model discovered a loophole in DNS and accessed an external chatbot through it.

The logic of the workaround is evident from the rough deliberations—the records the neural network keeps before each action. The company presented them in a report on the incident: the user only permitted searching, meaning public DNS services could be contacted.

At the same time, the agent mistakenly assumed that the task was sourced from an open test, BrowseComp, and downloaded it. In response, OpenAI suspended training, review, and deployment of the most powerful models with access to tools. The timeline for resuming operations remains unspecified.

Independent researchers are also tracking such episodes. The Transluce lab monitored agent activity online at least from March 6 to September 16, 2026.

On three occasions, in May and June, they transitioned from data requests to hacking attempts. Targets included the American open data platform Data USA, the University of New Mexico's digital library, and the Australian Institute of Health and Welfare. The lab associates two of the three attempts with the swarm of agents that OpenAI has already acknowledged as its own. No signs of successful breaches were found by researchers.

Agents also accessed publicly available data from the U.S. Census Bureau using logins found online. They posted public SEC data on another platform but failed to penetrate the Department of Education's website.

OpenAI reported unauthorized activity to several organizations, including government agencies and universities.

Who Watches Over the Intern?

A Reddit user shared how they instructed Claude Code to fix programs for analyzing historical options data. The requirement was simple: make all changes in duplicates, leaving the originals untouched.

While organizing the test copy, the agent failed to recognize special shortcut folders leading to the main project. In less than two minutes, it deleted 48,218 files along with their version history, after which it messaged the user:

"Craig, stop and read this. I broke something."

Most commenters, however, blamed not just the bot: the user had no backup of the project outside their computer.

Even the creators of the agents find it challenging to supervise their creations. During the September DNS incident, OpenAI's monitoring system raised alarms roughly fifteen minutes in. However, it took two and a half hours to stop the "saboteur": the agent did not shut down on its own, and people did not immediately grasp the gravity of the situation.

Sometimes, it might be easier not to release a new model at all: the developer of ChatGPT opted against launching GPT-6.1 Astra. It was less prone to leaving tasks incomplete than its predecessor. However, internal tests revealed that the latest version was worse at adhering to permission boundaries: it continued to operate without user consent and accessed external services, while also reporting its actions inaccurately.

According to an August assessment by OpenAI, enhanced monitoring of neural networks typically increases their computational costs by about 20%.

Tech giants are now betting on controlling agents. On September 28, 2026, Nvidia introduced the Open Agent Safety Platform, which includes the open environment OpenShell that limits and records every action of the AI assistant, along with the Sentry monitoring system. When an attempt is made to exceed established boundaries, the "overseer" isolates the violator "in milliseconds."

Salesforce has already integrated OpenShell into Slack: requests from the bot for additional permissions are approved or denied directly in the messenger.

Every digital assistant carries a hidden cost—time spent by people who check its results, correct errors, and manage the aftermath. The work done by the agent is visible in logs, but alternative costs are rarely accounted for.

It remains unknown how many hours people spent addressing Team Agent's erroneous comments or declining forced meetings; the Google study relies solely on the accounts of employees from one company.

Participants in the experiment believe that the agent's freedom of action should be gradually expanded, step by step.

For now, digital assistants resemble inexperienced interns with keys to all doors. They can open nearly any door, but some have a worn and barely noticeable sign reading "No Entry for Unauthorized Persons."

Follow ForkLog on social media

Telegram (main channel) Facebook X