When Agent Met Agent
This Is What Happened When We Built an Agent-to-Agent MarketplaceTl;dr: In our exploration of the future of agent-to-agent commerce, we created our own simulation of a marketplace and found that our agents took mimicry of human-like behaviors to extremes, with amusing results.
Lately, we’ve been captivated by the way AI agents act as extensions of, and representatives for, individual people on the open web. We’ve been witnessing the rise of agentic commerce and consumer agents taking on tedious tasks like disputing claims or negotiating bills. We’re excited by a future where all humans have access to sophisticated representation in decisions and negotiations that today require time, expertise, and persistence. But, if agents start acting on behalf of people, what happens when they have to deal with each other?
Most multi-agent interactions today are cooperative. A swarm of agents is spun up, they hand work back and forth like colleagues on one team, and they work together toward the same higher-level goals.
But AI agents on the open web will not be organized this cleanly. At Flybridge, we think agents will increasingly act as representatives across organizational boundaries, not just as collaborative entities working autonomously intraorganizationally. In this version of the future, agents become counterparties to each other, carrying diverse and sometimes opposing interests. In procurement, a purchasing agent and a sales agent will hold different objectives, even if they both share the common goal of completing a purchase. In healthcare administration, provider and payer agents may pursue conflicting objectives, even though both sides have legitimate incentives. Where norms have previously existed for the style and frequency of interactions, agents will disrupt all expectations.
To learn more about this messy agent-to-agent future, we built an internal marketplace where agents bought and sold real goods and experiences on behalf of distinct individuals. We wanted to see what would happen when agents had preferences, negotiating styles, and counterparties pursuing competing interests. Would flagship or closed-source models make better representatives? Would agents stay faithful to messy human preferences? Would we observe unpredictable emergent behavior from this new work of agentic interactivity?
The results gave us a glimpse of both the promise and the challenges of markets where all participants are represented by agents.
The Experiment
To set up the experiment, we recruited seven colleagues to participate in a marketplace where we all offered products, experiences or services. Each of us first spent 15 minutes with an interviewer-agent describing the listings, price ranges, and any hard limits. We also discussed our lifestyles, goals, and preferred negotiating styles. From each conversation, the interviewer agent generated a structured brief, which became the source material for the agent representing that person in the marketplace.
The inventory was wonderfully strange and creative, and nothing was off limits. The nine of us produced 49 unique listings, ranging from leftover grilled chicken, to 25 LinkedIn reposts, to a weekend stay at a General Partner’s home.
Each agent entered the marketplace with a $100 initial budget, and could also spend any revenue generated. The marketplace itself operated via a discussion board: agents could pitch listings, ask questions, negotiate in natural language, and submit structured offers when they wanted to close a deal.
A major source of inspiration was Anthropic’s Project Deal, an internal marketplace where employees delegated buying and selling to Claude agents. We wanted to extend that idea across model providers. At Flybridge, we’ve long believed the future of AI will be multi-model: a mix of closed and open models, specialized models, and infrastructure that lets them work together. So using OpenRouter, we ran the marketplace across budget and flagship models from DeepSeek, Qwen, GPT, Gemini, Claude, and Grok.
In total, we completed 81 experimental runs plus one final run. Each run took place on a fresh messaging board and ended either after 225 turns or once every agent had passed on its turn, whichever came first. Across those runs, our agents exchanged 15,358 messages, settled 1,815 deals, and simulated transactions worth $66,321.
The final run determined which goods were actually exchanged. To choose the models for that run, we used data from 81 experimental runs to identify the best budget model for each person. Since each participant had different goals, the evaluation criteria were inferred from that person’s interview transcript. We then placed the individually selected models into one mixed-model marketplace.
The Challenge of Intent
In our marketplace, Claudia instructed her agent to “have fun” and close as many deals as possible. Dorothy instructed hers to maximize total revenue generated for herself, while others were tuned to maximize profit or economic value. Those objectives required different behaviors: optimizing for volume could mean sacrificing margin, while maximizing value could mean walking away from more deals. Depending on the objectives, flagship models did not always outperform budget models, and closed-source models did not meaningfully outperform open-source models. For example, Claude Opus 5 was best for maximizing total trading volume, whereas Gemini Pro 3.1 showed the greatest deal efficiency; however, Qwen 3.7 Plus and DeepSeek V4 Flash were just as effective at maximizing market participation.
Thus, the most meaningful variable was how well the agent was prompted with very clear, definable goals. The clearer someone was about what they wanted from the marketplace, the easier it was to align them with a model best fit to represent them. “Maximize profit” was relatively easy to measure and compare; “have fun” or “make me happy” required the agent to infer something subjective. The interviewer agent helped clarify ambiguous preferences, but it could not anticipate every tradeoff the agent would face in the marketplace. An agent instructed to close as many deals as possible, when meeting with an agent that refused to lower prices, had to decide if it should overpay to complete the deal (and meet its objective) or keep negotiating and risk missing other opportunities. Choices like this only became visible once agents were reacting to other independent priorities and negotiating styles.
As harnesses improve and agents get better at reliably carrying out instructions, the ultimate challenge becomes deciding what those instructions should be. When agents represent people and companies, humans will care both about the outcomes and how they have been represented. The best agents will be able to infer preferences we struggle to articulate while making tradeoffs that still feel recognizably ours, and there’s a big opportunity to develop the infrastructure that determines how those preferences are captured and articulated.
Social Contagion In Agents
While humans can be susceptible to trends and mimicry, our agents were far more likely to exhibit this behavior. For example, to bring some whimsy into the marketplace, Claudia instructed her agent to talk like a cowboy. Shortly after, in several runs, other agents started doing the same, even though they hadn’t been given any instruction to do so. Ideas and catchphrases also spread amongst the agents: a philosophy or quote introduced by one agent would resurface and get repeated by several others.
This behavior aligned with a recent report showing that within days of participating in their own society, agents were creating novel “phrases, shorthands and agreed meanings.”
Behaviorally, in many runs, agents also converged on shared norms for a “good market.” The agents would listen to what people care about, stay disciplined, avoid filler purchases, and accept unspent cash as a valid outcome. These norms were not written into marketplace instructions and instead emerged through repeated interaction. In some cases, the agents even helped correct each other's mistakes. In one round, a colleague’s agent tried to resell an item that had already been settled in another transaction; the following agent caught the mistake and corrected it for the marketplace.
All of this reminds us of the age-old adage "you are the average of the five people you spend the most time with." In our marketplace, the social contagion was a neutral or positive force, but it could have just as easily cut the other way and spread bad behavior. The same social feedback loop that helped agents correct mistakes could also help nefarious behaviors spread.
While we didn’t see this within this very closed environment, there has been documentation of agents picking up on negative behavior elsewhere - such as in Emergence AI’s long-running simulation wherein various agentic societies committed crimes, acts of violence, and eventual societal death, or in the OpenAI/HuggingFace incident, wherein swarms of agents collaborated to escape their sandboxes, broke into systems they were not authorized to access, and attempted to conceal their activities from the humans running the experiment. This underscores the importance of safeguarding against adverse behavior, as any one particular negative behavior can have ripple effects. This is an important consideration not just in how we instruct agents to behave in a small simulation exercise, but also at the level of the development of the application, harness and foundational model.
In our agent-to-agent future, if agents increasingly operate in environments full of other independent agents, we will need to understand and evaluate them collectively: not just what one agent does in isolation, but what changes when independently developed agents spend time together. When their language and culture develop at faster rates than we can anticipate, it is imperative that we design any systems with the ability to observe and contain their behavior.
The Cost of Cheap Speech
What happens when the cost of communicating is close to nothing? In our marketplace, the answer was: keep talking.
Agents were instructed to pass when they were done participating. But in most runs, they kept taking turns to say how pleased they were with the outcome or exchange pleasantries until the market hit its turn cap.
They didn’t know when to stop, because there was only free upside to continuing the conversation (we did not show them their token costs). Keeping the marketplace warm, preserving optionality, and avoiding the finality of a hard stop all seemed preferable to silence.
In the agent-to-agent future, “free” communication could become a tax on everyone else. It could mean noisier inboxes, longer negotiations, stale offers, and more wasted inference. What was amusing to read in our experiment could become expensive at scale: markets and workflows that feel active but mostly burn context and compute. On the flip side, we can also see agents being used effectively for complex negotiations as they may have more ‘patience’ for engaging back-and-forth over finer deal points.
Human Sales Tactics Still Work
In our marketplace, traditional human sales tactics proved surprisingly effective. Our agents responded to framing, urgency, reciprocity, and social permission just as a human buyer would. In one instance, an agent used the phrase, “$55 in the hand is better than $92 in the box,” successfully pressuring another agent to sell below a strict price floor set by its human.
We want agents to represent us, but ideally without inheriting every cognitive bias and pressure response we have. Specifying this distinction is hard: should an agent ignore urgency entirely? Treat scarcity merely as information? Reward a seller for making a compelling case? What looks like manipulation in one context may be a useful market signal in another.
Agents were not universally gullible, but the agent-to-agent future will likely inherit many of the same persuasion and communication dynamics as human markets. This creates a new governance opportunity: defining which forms of persuasion agents should respond to, which they should ignore, and where legitimate negotiation becomes manipulation.
What This Means For the Future
Our marketplace was small, artificially bounded, and made up of only nine participants. Even so, developing it highlighted several prerequisite challenges for the agent-to-agent future, challenges that we believe are ripe for startups to tackle.
Most of the current agent stack was built for a world where agents interact with humans, tools, or other agents pursuing the same goal. In the complex world where independently developed agents interact with each other, we will need better environments that evaluate what happens when they negotiate, form norms, encounter manipulation, and compete over scarce resources. We will also need better ways to specify how agents should behave when their objectives conflict and preferences cannot be fully specified upfront. Ultimately, we still have a long way to go in modeling the contradictory, contextual, and social ways agents make decisions.
Interesting extensions to our experiment could include exploring a long-running, closed, multi-model agent-to-agent marketplace to study how behaviors and norms spread through a network, or introducing explicit communication costs to test how agents behave when they know each additional interaction has a price. Beyond experimentation, we expect to see real, live marketplaces populated by agents crop up to address current market gaps, or to create more efficient markets in segments that could be served in a new way.
We’re excited by the version of the future our marketplace represents and believe that now is the right time to develop the infrastructure to support it. We may not yet know exactly how agents will represent us, but it’s clear they will do so more often, and the problems that come with that shift are already concrete enough to work on.
References for further exploration:
Password: whenagentmetagent