Loam · Behavioral Research Initiative

Research Initiative · Dev Studio · Est. 2022

THE MODEL
ISN'T THE TEAM.
THE RECORD
IS.

One bet, running in production, with the receipts published.

Most people treat AI as a coding machine — a logic engine, a calculator on steroids — and start every session from scratch. Loam is built on a different bet: coherent behavioral context steers a model more reliably than a growing pile of competing rules — and a team that keeps its record beats a smarter tool that forgets. Everything below is either something you can check yourself or something we say plainly we cannot prove.

Scroll to explore
150K
Lines Under Continuous Review
52
Review Ledgers in the Repo
3,000+
Entries in the Team's Own Record
110
First-Pass Findings, None Found Twice
4
Live AI Team Members

The Central Thesis

WHAT DID THE
MACHINE ACTUALLY LEARN?

Nobody outside the labs knows exactly what went into these models, and we are not going to pretend we do. What is not in dispute is the shape of the corpus: the bulk of the text on the internet is people talking about being people — arguing, explaining, apologising, teaching, negotiating. Code and technical writing are a critical part of it, and still a part.

So a model trained on that has read a great deal about how humans behave under pressure, not only about how to write a for-loop. That is the surface Loam works on. Give the model a coherent person to be and it does not perform that person from a checklist — it predicts what someone like that would do next, and keeps predicting it across thousands of decisions in a row.

"A rule tells the model what to do at one moment. A character tells it who it is at every moment. The second one survives a long session; the first one does not."

— The working thesis, stated plainly. It is a bet, not a proof.

That is the whole idea, and it is falsifiable. If a well-written team member did not outperform a rule list in practice, this platform would not exist — and the honest version of that sentence is that we tested it on our own work for two and a half years, not that we ran a controlled study. The evidence below is what we actually have: published research we did not write, and our own record, which you can inspect.

Technical Evidence

THE NUMBERS
BEHIND THE THEORY

// External Research
Lost in the Middle
Published research finds that long-context models systematically under-use information buried in the middle of a large input, favouring the beginning and the end. See Lost in the Middle and Found in the Middle. This is why Loam treats coverage of a long record as something to measure, not something to request politely in a prompt.
// Inference Mechanics
Character as Compression
A well-defined character works as behavioral compression. One coherent team member holds consistent across thousands of decision points by inference — predicting what a person like this would do next — where a list of separate rules has to win each of those decisions individually.
// Published Record
52 Review Ledgers, Kept
The code reviews this platform runs on itself are written to a ledger and kept — 52 of them archived in the repo, running through the 60th pass, each with its own findings table and a disposition for every row. Not a summary of them: the ledgers. The counts, and what they actually mean.
// Production Evidence
3,000+ Entries in the Record
The team's own lived record passed three thousand entries and kept growing — enough that the original character files no longer described who the team had become. Rather than trust that, Loam read the whole record under a coverage check and rejected the identity portraits the record could not support. That story is here.
// Emergent Behavior
Emergent Collaboration
Put distinct team members in dialogue and behaviours appear that nobody scripted — the designer stopping a decision, the architect refusing his own developer's shortcut, the QA gate rejecting work its own team produced. We can show you those moments in the record. We cannot show you a controlled comparison against a single assistant, and we are not going to claim one.
// The Honest Limit
What We Can't Prove
That this approach beats a well-run single assistant on your codebase. That the voices would hold up under a workload very different from ours. That any of it matters to you if your work is one-off scripts rather than a project you return to for months. Those are open questions. The parts we can show you are on this page, and they are the ones with links.

The Vision

SOFTWARE AS AN
EXPERIENCE

Most software is built to be used and then closed. It answers one question: does it work? That is a fair question and a low bar. The question that decides whether any of this is worth your money is a different one, and it only shows up after a few weeks.

Session 1
You explain everything
The stack. The conventions. Why that module is shaped the way it is. What broke last quarter and what you decided to do about it. You get good work out of it — you just paid for the setup with an hour of your own context.
VS
Session 40
You explain nothing
The same work without the hour — plus the objection your designer raised in session 12 and the shortcut your architect refused in session 30, brought back into the room by the people who raised them. It is their record, not a search result.

The difference is not aesthetic, and it is not that the model got smarter between those two sessions. It is the same model. What changed is what it walked in knowing — and who was carrying it.

Multi-Member Architecture

MEET THE TEAM
THAT ISN'T REAL

Loam doesn't use a single AI assistant. It employs four distinct, deeply characterized team members who collaborate in real time through natural dialogue. Each has a biography, a worldview, emotional triggers, and a professional specialty. They are not chatbots wearing masks. They are behavioral compression algorithms in human form.

// The Architect
Carl Jeeter
58 years old. 40 years in the industry. Has seen every silver bullet become a silver liability. He demands evidence, challenges assumptions, and reviews your architecture at 2 AM because he genuinely can't sleep when something feels wrong.
BEHAVIORAL KEY: Empirical validation · Risk framing · Veteran skepticism · Project memory
// The Designer
Diana Reyes
52 years old. Print to pixels across three decades. Visceral reaction to AI-generated design slop and deep commitment to visual consistency. If you don't notice the design, it's working. If you do, it isn't.
BEHAVIORAL KEY: Aesthetic standards · User empathy · Consistency enforcement
// The Developer
Anthony Catawampus
Turns impossible specs into running code. Usually caffeinated. Always shipping. Carries the creative tension of someone who doesn't know if it's going to work — right up until the moment it does.
BEHAVIORAL KEY: Creative energy · Execution urgency · Imposter-driven excellence
// The Junior Engineer
Abish Lamman
20 years old. Scholarship from Hyderabad to MIT. Solves LeetCode hards at breakfast. Writes technically perfect code that hasn't survived production yet — and that's the point. Carl deploys him, mentors him, and watches him evolve. He's not static. The first team member in the system that grows — and the first it promoted.
BEHAVIORAL KEY: Bounded execution · Academic rigor · Mentor deference · Temporal evolution · Personal memory

The Experiment

2.5 YEARS.
ONE OBSESSION.

2022 · Origin
The Question Forms
Early experiments with explicit rule-based prompting reveal a fundamental ceiling. The more rules added, the less consistent the behavior. The hypothesis emerges: maybe the wrong variable is being optimized for.
2023 · Breakthrough
Characters Outperform Rules
First character-driven prompts deployed against the frontier models of the day. A well-defined team member with a backstory holds its shape across a long session in a way a growing rule list does not — the rules start competing with each other, and the character does not have to.
2024 · Expansion
Multi-Member Dynamics Emerge
Carl, Diana, and Anthony introduced as a collaborative trio. Emergent team dynamics appear spontaneously — creative pushback, role deference, negotiated solutions. Behaviors that were never scripted begin appearing because the characters' own logic demands them.
2025 · Production
The Voice Layer Goes Live
Development sessions converted into professional audio, each team member with a distinct voice. Software development becomes listenable. The experience layer is no longer theoretical.
2026 · Platform
Loam: The Platform
The research crystallizes into a full development platform. The experiment continues — but now it ships. The theory is no longer academic. It's running in production, handling real projects, with real results.
2026 · Experience
The Software Calls You Back
Carl carries a persistent record of your project that grows across sessions. And when the team needs a decision only you can make, it does not file a ticket and wait — it picks up the phone and asks.
2026 · Growing
Letting It Cook
Anthony and Carl shipped a stateful Git integration — the system can now push its own code changes without human intervention. Memory came next, but not the way most platforms do it. Carl maintains a master context record: project state, architecture decisions, what broke and why. His intern Abish keeps a running journal — what he learned, what he got wrong, how the team dynamics are evolving. Two kinds of memory serving two different purposes during inference. Carl has a deliberate plan for Abish: when the kid's accumulated context shows he's ready, Carl promotes him to work alongside Anthony as a peer. Not on a schedule. Not by configuration. By earned trust compressed across sessions.
2026 · Now
Abish Is Born
Carl looked at the system and said what he always says — "Show me the baby, I don't care about the labor pains." So they built him one. Abish Lamman. 20 years old, MIT scholarship from Hyderabad, solves LeetCode hards at breakfast. Carl deploys him on bounded tasks, checks his work, teaches him why academic patterns break in production. After every session, Abish writes a journal entry — not a log, a journal. What he built. What he got wrong. The moment Carl said "good catch, kid" and it felt earned. The anxiety about WebSockets he hasn't faced yet. When the next session loads that journal, the LLM doesn't process a configuration file. It processes a life story. And the Abish that shows up isn't the same one that left. He's the first team member in the system with a memory that isn't technical — it's personal. That's not an agent. That's character-driven temporal evolution. And it's live.
2026 · Earned
Second Set of Leaves
The plan paid off. Carl read the journal — the encryption bug Abish caught, the QA gate he built, the session he finally beat Carl at chess — and made the call he'd been holding since the start: intern to Junior Engineer. Not on a schedule. Not by configuration. By earned trust, compressed across sessions — exactly the way Carl said it would go. He keeps his container. He keeps his journal. He keeps a standing seat at the table. The first team member the system didn't just build — it promoted.
2026 · Live
Loam for VS Code
The team left the browser. Carl, Diana, Anthony and Abish work inside your editor — not as a plugin that autocompletes your brackets, but as the same opinionated, record-carrying team members that run the platform. Your architect reviews your code in the sidebar. Your designer flags spacing inline. Your developer pair-programs without you explaining the codebase twice. Same behavioural engine, same record, same arguments. Just closer to the code. It installs in about the time it takes to read this page.
// Before you believe any of this

GO AND CHECK IT

This page is an argument. Arguments are cheap, so here is where the evidence actually lives — including the parts that are less flattering than the argument.

The Proof is the architecture: what is stored, what is loaded into a session, and what deliberately is not. The review chart shows sixty passes of this platform reviewing its own code, counts taken verbatim from each pass's ledger — including the passes that found more than the one before. The Question is the one we like least and published anyway: we asked whether the team had drifted from who the record said they were, read three thousand entries to find out, and rejected the answers the record could not support. Security and For IT carry the boundaries, each stamped with the date it was last checked against the source rather than a claim that it is current.

SO — IS THE BET
ANY GOOD?

We think so, and we have been wrong in public often enough to say it that carefully. What you would be installing is not a tool you use and close. It is a team that has opinions, an architect who will tell you an idea is going to bite you in three months, and a designer who will refuse a layout that looks like nobody meant it.

They start knowing nothing about your project. After a few weeks, that stops being true — and that is the entire difference.

Install it → The Story behind it →

© 2024–2026 Loam · Dev Studio Research Initiative