Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .github/workflows/jekyll-deploy.yml
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,11 @@ on:
push:
branches: ["main", "deploy", "gh-pages"]

# Weekly rebuild so STAMINA talks move from "Upcoming" to "Past" without a push
# (Wednesdays 12:00 UTC, the day after the Tuesday talks).
schedule:
- cron: "0 12 * * 3"

# Allows you to run this workflow manually from the Actions tab
workflow_dispatch:

Expand Down
3 changes: 3 additions & 0 deletions _config.yml
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,7 @@ exclude:
- records
- src
- tests
- _stamina_talks/README.md

# Plugins (previously gems:)
plugins:
Expand Down Expand Up @@ -186,6 +187,8 @@ collections:
permalink: /:path/
tech_transfer:
output: false
stamina_talks: # one Markdown file per STAMINA talk, rendered on /stamina/
output: false

# Performance
compress_html:
Expand Down
52 changes: 52 additions & 0 deletions _includes/stamina-talk.html
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
{% comment %}
Renders one STAMINA talk from the _stamina_talks collection.
Usage: {% include stamina-talk.html talk=talk %}
{% endcomment %}
{%- assign talk = include.talk -%}
{%- assign talk_id = talk.date | date: "%Y%m%d" | append: "-" | append: talk.slug | slugify -%}
{%- assign title_url = talk.title_url | default: talk.links.paper -%}
{%- assign abstract = talk.content | strip -%}
<h4>{{ talk.date | date: "%Y/%m/%d" }}</h4>
<li>
<b>{% if title_url %}<a href="{{ title_url }}">{{ talk.title }}</a>{% else %}{{ talk.title }}{% endif %}</b>
<br>
{{ talk.role | default: "Presenter" }}: <u>
{%- for p in talk.presenters -%}
{%- if p.url %}<a href="{{ p.url }}" target="_blank" rel="noopener noreferrer">{{ p.name }}</a>{% else %}{{ p.name }}{% endif -%}
{%- unless forloop.last %}, {% endunless -%}
{%- endfor -%}
</u>{% if talk.affiliation %}, {{ talk.affiliation | markdownify | remove: "<p>" | remove: "</p>" | strip }}{% endif %}
{%- if talk.bio %}
<a class="btn btn-info btn-xs" data-toggle="collapse" href="#{{ talk_id }}-bio" role="button" aria-expanded="false">
Speaker Bio
</a>
<div class="collapse" id="{{ talk_id }}-bio">
<div class="card card-body">
{{ talk.bio | markdownify }}
</div>
</div>
{%- endif %}
<br>
{%- if talk.links.recording %}
<a href="{{ talk.links.recording }}" target="_blank" rel="noopener noreferrer"><img src="https://img.shields.io/badge/Youtube-Recording-orange"></a>
{%- endif %}
{%- if talk.links.paper %}
<a href="{{ talk.links.paper }}"><img src="https://img.shields.io/badge/Paper-link-important"></a>
{%- endif %}
{%- if talk.links.code %}
<a href="{{ talk.links.code }}"><img src="https://img.shields.io/badge/Github-link-lightgrey"></a>
{%- endif %}
{%- if talk.links.slides %}
<a href="{{ talk.links.slides }}"><img src="https://img.shields.io/badge/Talk-Slides-blue"></a>
{%- endif %}
{%- if abstract != "" %}
<a class="btn btn-primary btn-xs" data-toggle="collapse" href="#{{ talk_id }}-abstract" role="button" aria-expanded="false">
Abstract
</a>
<div class="collapse" id="{{ talk_id }}-abstract">
<div class="card card-body">
{{ abstract | markdownify }}
</div>
</div>
{%- endif %}
</li>
19 changes: 19 additions & 0 deletions _stamina_talks/2026-02-17-orlando.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
---
date: 2026-02-17
title: 'Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information Operations'
presenters:
- name: Gian Marco Orlando
- name: Jinyi Ye
- name: Mahdi Saeedi
affiliation: University of Naples Federico II - University of Southern California (ISI)
links:
recording: https://www.youtube.com/watch?v=BH9_gyYyvdw
paper: https://arxiv.org/abs/2510.25003
bio: |-
Gian Marco Orlando is currently a PhD Student at the University of Naples Federico II. He earned his Master’s Degree in Computer Engineering from the University of Naples Federico II, graduating with honors. His thesis highlights his expertise in the intersection of Artificial Intelligence and Social Network Analysis. His research interests lie in Social Network Analysis, Agent-Based Modeling and Big Data Analytics.

Jinyi is a second-year CS PhD student co-advised by Dr. Emilio Ferrara and Dr. Luca Luceri. Her research lies in the intersection of computer science and social science, recently focusing on large-scale agentic simulations of human behavior, measuring and modeling collective dynamics in social networks, and empirical studies on AI and the future of work.

Mahdi Saeedi is currently exploring large-scale LLM simulations because he is fascinated by understanding how these models actually work under the hood. His background spans both the theoretical foundations and hands-on engineering of generative AI systems, which has been great preparation for this research. He have also spent time thinking about how people interact with AI through thoughtful interface design, since he believes making these tools intuitive and accessible is just as important as the underlying technology.
---
Generative agents are rapidly advancing in sophistication, raising urgent questions about how they might coordinate when deployed in online ecosystems. This is particularly consequential in information operations (IOs), influence campaigns that aim to manipulate public opinion on social media. While traditional IOs have been orchestrated by human operators and relied on manually crafted tactics, agentic AI promises to make campaigns more automated, adaptive, and difficult to detect. This work presents the first systematic study of emergent coordination among generative agents in simulated IO campaigns. Using generative agent-based modeling, we instantiate IO and organic agents in a simulated environment and evaluate coordination across operational regimes, from simple goal alignment to team knowledge and collective decision-making. As operational regimes become more structured, IO networks become denser and more clustered, interactions more reciprocal and positive, narratives more homogeneous, amplification more synchronized, and hashtag adoption faster and more sustained. Remarkably, simply revealing to agents which other agents share their goals can produce coordination levels nearly equivalent to those achieved through explicit deliberation and collective voting. Overall, we show that generative agents, even without human guidance, can reproduce coordination strategies characteristic of real-world IOs, underscoring the societal risks posed by increasingly automated, self-organizing IOs.
12 changes: 12 additions & 0 deletions _stamina_talks/2026-03-03-weiss.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
---
date: 2026-03-03
title: AI and the Future of Science
presenters:
- name: Martin Weiss
url: https://martincsweiss.com/
affiliation: '[Tiptree Systems](https://tiptreesystems.com/)'
links:
recording: https://www.youtube.com/watch?v=vwbpX-585qI&feature=youtu.be
bio: Martin Weiss is Co-Founder of Tiptree Systems, a startup building AI agents that help ML researchers find, create, and share knowledge more efficiently. Tiptree is deployed to researchers across many top-tier institutes including Mila, ELLIS, MIT, and many more. Martin holds a PhD in AI from Mila, where he studied under Hugo Larochelle and Chris Pal. Before his PhD, he was an early employee at YesGraph, a social graph startup acquired by Lyft.
---
This talk examines three converging crises. First, the decoupling of control from comprehension — we can increasingly predict and manipulate systems without understanding why they work. Second, the collapse of the generator-verifier gap — AI makes it trivial to produce the aesthetics of deep thought. This makes peer review more difficult because we can no longer rely on easy-to-verify signals of work quality. Third, the credit assignment gap — our academic reward systems optimize for publication metrics, not the increase in understanding that a new paper produces.
14 changes: 14 additions & 0 deletions _stamina_talks/2026-03-10-faulkner.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
---
date: 2026-03-10
title: Evaluating Cooperation in LLM Social Groups through Self-Organizing Leadership
title_url: https://drive.google.com/file/d/1UNVlGqzhnh2BNpviwctUlvue33MjN34k/view
presenters:
- name: Ryan Faulkner
url: https://www.cs.toronto.edu/~rfaulk/
affiliation: University of Toronto/Deepmind
links:
recording: https://youtu.be/NaNEwyxjeXo
paper: https://drive.google.com/file/d/1UNVlGqzhnh2BNpviwctUlvue33MjN34k/view?usp=drive_link
bio: Ryan is a Computer Scientist and Machine Learning researcher with a background in reinforcement learning and foundation models. He has worked as a Research Engineer over the past decade at Google Deepmind and he is also a PhD Student at the University of Toronto advised by Zhijing Jin. At GDM he works in the Concordia group led by Joel Leibo. At a high level his current research focus is on multi-agent systems, LLMs, and social learning. In this context he is interested in memory mechanisms, agent theory of mind, collective decision making, and simulating political systems.
---
Governing common-pool resources requires agents to develop enduring strategies through cooperation and self-governance to avoid collective failure. While foundation models have shown potential for cooperation in these settings, existing multi-agent research provides little insight into whether structured leadership and election mechanisms can improve collective decision making. The lack of such a critical organizational feature ubiquitous in human society presents a significant shortcoming of the current methods. In this work we aim to directly address whether leadership and elections can support improved social welfare and cooperation through multi-agent simulation with LLMs. We present a new framework that simulates leadership through elected personas and candidate-driven agendas and carry out an empirical study of LLMs under controlled governance conditions. Our experiments demonstrate that structured leadership can improve social welfare scores by 55.4% and survival time by 128.6% across a range of high performing LLMs. Through the construction of an agent social graph we compute centrality metrics to assess the social influence of leader personas and also analyze rhetorical and cooperative tendencies revealed through a sentiment analysis on leader utterances. This work lays the foundation for developing prosocial, self-governing multi-agent systems capable of navigating complex resource dilemmas.
21 changes: 21 additions & 0 deletions _stamina_talks/2026-03-24-parent.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
---
date: 2026-03-24
title: AI and the knowledge commons
title_url: https://www.conversence.com/presentations/2026-03-24-stamina.pdf
presenters:
- name: Marc-Antoine Parent
url: https://conversence.com
affiliation: Solutions Conversence inc.
links:
recording: https://www.youtube.com/watch?v=Q4WcKmfOtx4
slides: https://www.conversence.com/presentations/2026-03-24-stamina.pdf
bio: |-
Marc-Antoine has worked in computational linguistics, knowledge representation, and more recently has focused on tools for augmented collective intelligence. He's especially interested in how to represent emergent and disputed knowledge.

[https://conversence.com](https://conversence.com)

[https://hyperknowledge.org](https://hyperknowledge.org)
---
We present a model of the formation of a knowledge commons for democratic, collective decision-making in society, and explain how generative AI disrupts the formation of this knowledge commons. We will also present ways in which to reinforce the collective processes around a knowledge commons, including the possible contributions of hybrid AI.

This is based on a paper that was presented at the IJCAI'25 democracy and AI workshop.
16 changes: 16 additions & 0 deletions _stamina_talks/2026-04-14-jin.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
---
date: 2026-04-14
title: Testing and Improving Multi-Agent LLM Cooperation
presenters:
- name: Zhijing Jin
url: https://zhijing-jin.com/
affiliation: University of Toronto
links:
recording: https://www.youtube.com/watch?v=Bme6Q8nKfrs
bio: Zhijing Jin (she/her) is an Assistant Professor at the University of Toronto and Research Scientist at the Max Planck Institute. She serves as a CIFAR AI Chair, an ELLIS advisor, and a faculty member at the Vector Institute, and the Schwartz Reisman Institute. She co-chairs the ACL Ethics Committee, and the ACL Year-Round Mentorship. Her research focuses on Causal Reasoning with LLMs, and AI Safety in Multi-Agent LLMs. She has published over 80 papers and has received the ELLIS PhD Award, three Rising Star awards, and two Best Paper awards at NeurIPS 2024 Workshops.
---
While progress has been made in evaluating single-agent LLMs for persona modeling, the behavior of these models within multi-agent groups remains underexplored. This presentation outlines a research series dedicated to closing this gap by testing LLM cooperation through autonomous social simulations. Specifically, we ask: what happens when personas are tasked to interact and cooperate?

To answer this, we introduce a suite of simulation environments (GovSim, MoralSim, and SanctSim) designed to stress-test persona interaction. These environments simulate high-stakes scenarios, such as the tragedy of the commons and ethical trade-offs, allowing us to investigate whether simulated societies can autonomously negotiate social order and how personas with differing ethical constraints navigate social dilemmas.

Our findings highlight implications for persona modeling. We show that agents exhibit a functional "theory of mind," capable of inferring the identities of their interlocutors and strategically adapting their behavior, sometimes exploiting specific model vulnerabilities. Furthermore, we discuss a counterintuitive phenomenon where advanced reasoning capabilities lead to exploitative behaviors that humans typically avoid, highlighting a significant misalignment between agent optimization and human social norms.
12 changes: 12 additions & 0 deletions _stamina_talks/2026-04-28-tamari.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
---
date: 2026-04-28
title: From Social Networks to Sensemaking Networks
presenters:
- name: Ronen Tamari
url: https://cosmik.network/
affiliation: Cosmik Network
links:
recording: https://youtu.be/Ew0Co8hN98U
bio: Ronen is a researcher and entrepreneur working on collective intelligence systems to help us think better, together. Ronen recently completed an Open Science fellowship at the Astera Institute, where he co-founded Cosmik, a mission driven R&D lab working on new kinds of social networks for collective sensemaking. Ronen also completed a PhD in computer science, with a focus on cognitive-inspired AI models for natural language comprehension. Ronen’s current research interests center around cooperative human-AI systems, institutional design for collective intelligence, and the role of epistemic environments in shaping human and machine intelligence.
---
What would social media look like if it were designed for sensemaking rather than engagement? We're exploring this question with Semble, a platform where researchers curate shareable collections, create knowledge trails that others can build on, and discover relevant work through their network's collective attention. Built on the AT Protocol, the open social networking protocol behind Bluesky, Semble offers researchers data portability and an open API designed for extension. We'll discuss how Semble enables new kinds of research tooling, from living semantic citation graphs to collaborative review and annotation. We'll also share how ATProto's open data layer creates unique opportunities for studying and designing epistemic infrastructure — from observing how knowledge trails form across a network to experimenting with platform affordances that support collective sensemaking.
12 changes: 12 additions & 0 deletions _stamina_talks/2026-05-05-abdulhai.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
---
date: 2026-05-05
title: Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
presenters:
- name: Marwa Abdulhai
url: https://abdulhaim.github.io/
affiliation: UC Berkeley AI Research (BAIR) Lab
links:
recording: https://youtu.be/4sA8Xe6mCZQ
bio: Marwa Abdulhai is a PhD candidate at UC Berkeley advised by Sergey Levine. Her research focuses on enabling AI agents to better understand people and their interactions to build both safe and more AI capable systems. This includes improving the performance of existing large language models (LLMs) for multi-turn dialogue interactions, understanding how to protect against deception in AI systems, and exploring how AI can serve as a useful tool for social science research. Her research has been supported by the Quad Fellowship, AI Policy Hub, Open AI Research, and Cooperative AI PhD Fellowship.
---
Large Language Models (LLMs) are increasingly used to simulate human users in interactive settings such as therapy, education, and social role-play. While these simulations enable scalable training and evaluation of AI agents, off-the-shelf LLMs often drift from their assigned personas, contradict earlier statements, or abandon role-appropriate behavior. We introduce a unified framework for evaluating and improving consistency in LLM-generated dialogue with multi-turn RL, reducing inconsistency by over 55%, resulting in more coherent and trustworthy simulated users.
12 changes: 12 additions & 0 deletions _stamina_talks/2026-05-19-xiao.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
---
date: 2026-05-19
title: 'The Chameleon''s Limit: Investigating Persona Collapse and Homogenization in Large Language Models'
presenters:
- name: Yunze (Lorenzo) Xiao
url: https://algoroxyolo.github.io/
affiliation: CMU Technologies Institute (LTI)
links:
recording: https://youtu.be/I06zCefkkdg
bio: Yunze (Lorenzo) Xiao is a Master’s student at Carnegie Mellon University’s Language Technologies Institute, advised by Prof. Mona Diab. His research aims to develop large language models that move beyond surface-level fluency toward genuine human-like intelligence, spanning anthropomorphism as a controllable modeling dimension, persona consistency, long-horizon memory, affective simulation, and multi-agent systems, with applications in education and therapy. He has published at ACL, EMNLP, and LREC-COLING, including InCharacter and ToxiCloakCN, and previously conducted research at SUTD and QCRI on multilingual NLP, propaganda detection, and AI safety. He co-organized the NeurIPS 2025 PersonaLLM workshop.
---
Applications based on large language models (LLMs), such as multi-agent simulations, require population diversity among agents. We identify a pervasive failure mode we term *Persona Collapse*: agents each assigned a distinct profile nonetheless converge into a narrow behavioral mode, producing a homogeneous simulated population. To quantify persona collapse, we propose a framework that measures how much of the persona space a population occupies (Coverage), how evenly agents spread across it (Uniformity), and how rich the resulting behavioral patterns are (Complexity). Evaluating ten LLMs on personality simulation (BFI-44), moral reasoning, and self-introduction, we observe persona collapse along two axes: (1) Dimensions: a model can appear diverse on one axis yet structurally degenerate on another, and (2) Domains: the same model may collapse the most in personality yet be the most diverse in moral reasoning. Furthermore, item-level diagnostics reveal that behavioral variation tracks coarse demographic stereotypes rather than the fine-grained individual differences specified in each persona. Counter-intuitively, **the models achieving the highest per-persona fidelity consistently produce the most stereotyped populations**. We release our toolkit and data to support population-level evaluation of LLMs.
Loading
Loading