Perpetual Agents: Beyond the Task
What changes when an AI agent takes on an ongoing responsibility instead of a task with an endpoint.
There’s something magical about giving an AI agent a project, going to bed, and waking up to a working prototype. I’ve had that experience building a legislative bill tracker for policy analysts. By morning, I had something I could open and try. But what’s stayed with me more is watching a project keep moving when I’m busy with something else.
I’ve been building a personal productivity app called Day Desk. Along the way, I gave a coordinating agent in Codex an ongoing job: keep developing and maintaining the software. Over several weeks, it managed the development backlog, assigned work to coding agents, checked their results, and shipped improvements to Day Desk. I received periodic summaries of what it completed and could change priorities or suggest something new, without having to return each morning and supply the next instruction.
It felt a little like having my own team. I hadn’t expected to feel that way about software. Having been a CEO, I recognized something familiar in the experience: I still wanted to know what was happening, and I still had opinions about the details, but I was spending more time setting direction and reviewing results. I was getting used to reading an update instead of being the person who got everything moving.
That experience helped me put a name to what interested me: a perpetual agent, with a responsibility that continues beyond the completion of any individual task.
“Build a bill tracker” has an endpoint, even if I haven’t worked out exactly what finished looks like. “Maintain and improve this system” keeps creating work as needs change and problems emerge. When something ships, the coordinator looks through the backlog and figures out what to work on next. Shipping is a milestone within the responsibility, rather than the end of it.
This is where I find it useful to distinguish persistent from perpetual. A persistent agent can retain context and resume work across sessions or interruptions. It might spend weeks building an application, remembering decisions and recovering from setbacks, and eventually finish. A perpetual agent needs that continuity too, but its mandate survives the release. Customers encounter problems, dependencies change, and new capabilities become possible. It needs to keep asking whether the system is serving its purpose and what deserves attention next.
Persistence describes continuity of the agent. Perpetuity describes continuity of its responsibility.
The terms overlap in practice. I’m using this distinction to describe how we assign and manage the work, rather than claiming an entirely new category of technology. It also helps explain what I mean by “always on.” The agent doesn’t need to run continuously. It can wait for an issue, a change in conditions, or a scheduled review. Its responsibility remains in place until a person changes or ends it.
Once I started thinking about Day Desk this way, maintaining a backlog no longer seemed like the whole job. I wanted the coordinator to help reconsider the roadmap itself: what was working, where quality was falling short, and what might be worth improving next. To do that well, it would also need to check in with me about how my needs were changing.
In my setup, periodic check-ins bring the coordinator back to review the project. Markdown files capture requirements, decisions, and progress between sessions; summaries help carry context forward; and Linear tracks the work. Together, those give the coordinator a way to pick up where it left off and decide which workers should handle what comes next.
I still steer the project, and some problems need my involvement. An ongoing responsibility makes the boundaries more important: what the agent can change, what it can spend, and when it needs my review. I also don’t want it inventing work just to stay busy. Sometimes the right decision is to wait, and a system responsible for the long term needs to recognize that.
The surprise of an overnight prototype is immediate. With Day Desk, the appreciation came more gradually. I’d check in, read what shipped, and take a moment to notice how far it had come. It was a little like the feeling when a team starts handling things you used to have to stay on top of.
I still care about the details. I can get involved whenever I want. But I’m spending more time thinking about what the system should become and less time telling it what to do next.
Written with drafting and editing assistance from Codex. The experiences and final editorial decisions are my own.
I work at OpenAI. The views expressed here are my own and do not necessarily reflect those of my employer.