You run a portfolio of small apps with AI agents by giving each agent one narrow job, keeping the keys for each app separate, and approving anything that sends, pays, deletes or publishes yourself. Then you spend one fixed block each week checking every app for the same few failures. It is boring on purpose, because several apps break at the same time when one agent can touch all of them.
What this post is, and what it is not
I operate Riot Digital, an AI automation agency, and I've shipped three small products: Outposts, Columns and DrawDog. This post is a framework built from published guidance and my own judgment. It is not a case study.
I'm not giving you revenue, user counts or hours saved. I don't have numbers I can back up, and a post full of them would be hype. Every item below carries one of three labels:
- Sourced: published guidance, linked. Most of it comes from vendor blogs, so read it as the publisher's view.
- My rule: a judgment call I'm recommending. Not a measured result.
- Speculative: an idea worth knowing about that you should not bet on yet.
Step 1: Give each agent one job
Label: sourced. The pattern across the pages that rank for this topic is simple. The founder sets direction and approves decisions with financial, legal or reputational impact. Each agent owns one scoped workflow and stops at an approval step. Alook describes this model, and it sells an agent product, so keep that in mind.
Typical roles are a coding agent, a marketing agent, a support agent for tickets and an analytics agent that watches your key numbers. Treat that list as a starting menu, not a tested setup.
For a portfolio, add one more rule. Label: my rule. An agent works on one app at a time. If you want the same kind of help on three apps, you run three separate agents, not one agent with three sets of keys.
Step 2: Draw an access map
Label: sourced, then my rule. Least privilege means an agent gets only the tools, files, credentials and network access its current task needs, and that access ends when the task ends. SpecStory defines it this way. The same idea shows up in Auth0's write-up of the OWASP agent risk list: a code analysis task should not be handed database write or file deletion tools.
Why it matters more for a portfolio: Obsidian Security notes that an agent connected to several systems exposes the combined authority of every permission it holds. One agent holding keys to all your apps is the worst version of that.
Make a one page table. One row per agent. Here is a made-up example so you can see the shape:
| Agent (made-up example) | Can read | Can change | Keys it holds | Needs my yes for |
|---|---|---|---|---|
| Code agent, app A | App A repo | App A branch, not main | App A repo only | Merging to main, deploys |
| Support agent, app B | App B inbox | Draft replies | App B inbox only | Sending any reply, refunds |
| Marketing agent, app C | App C analytics | Draft posts | None | Publishing anything |
If you can't fill in the "keys it holds" column for an agent, you don't know your risk yet. Fix that first.
Step 3: Make a list of things that always wait for you
Label: sourced, with my rule on the list itself. Security guidance from Graylog and Auth0 says to require human approval for irreversible actions such as sending external mail, moving money or deleting records. An agent that is allowed to delete should need approval for each deletion.
My rule for the list is short. Anything that sends, pays, deletes or publishes waits for you. Start there. Loosen it one action at a time, only after you've watched the agent handle that action properly for a while. Never loosen the pay and delete rows.
Step 4: Run a 20 minute weekly check
Label: my rule. Pick one fixed slot, ideally the same one every week. If you have a day job, a weekend morning works. For each app, ask four questions and write one line per answer:
- Did it run? Look for an error or a gap in what the agent should have produced.
- Did anything change that I did not approve?
- Is there one decision this app is waiting on from me?
- Is this app still the thing I meant it to be?
Four questions across a few apps fits in about 20 minutes only if you keep the notes short. If it takes longer, you probably have too many apps with live agents, or the agents are doing too much. That is useful information.
Step 5: Know what breaks
Here is the short list of failures to watch for, each labeled.
Scope creep. Label: sourced, one person's account. A DEV Community writer who let an AI agent run a SaaS reports that it quietly grew one codebase into five products, adding monitoring features, calculators and lead capture. Each addition looked reasonable alone. The missing piece, the writer says, was someone to review the growing list and decide what the product actually is. That is question 4 in your weekly check. It is one story, not a study.
Silent failures. Label: my rule. An agent that fails quietly looks the same as an agent that has nothing to do. Ask each agent to leave one line behind every run, even when nothing happened. No line means look.
Drift and stale context. Label: speculative. Some write-ups describe agents losing track of earlier instructions as their working context grows, and a few propose engineering fixes. I can't tie that to one source I'd stand behind, so treat it as a thing to watch for, not a settled fact. The cheap fix is the same one: short tasks, one app, a clear start each time.
No single owner for approvals. Label: my rule. If two agents can both approve each other's work, nobody is approving. You are the gate.
Prompt injection, in plain words
Agents read text: web pages, tickets, emails, comments. Prompt injection is when that text contains instructions the agent should not follow, and the agent follows them anyway. Anyone can put text in front of your agent, so you can't count on the agent to tell the difference.
Label: sourced. Obsidian Security says least privilege reduces the damage of a compromise but does not prevent prompt injection, and that no current defense fully does. That is the honest summary. You are shrinking the blast radius, not building a wall.
That is also why the access map comes first. If a support agent reads a hostile ticket and can only draft replies for app B, the worst case is a bad draft you will read before it goes out. If it holds keys to every app, the worst case is much larger.
Label: speculative for a solo builder. Auth0's guidance mentions short-lived credentials issued for one task, plus audit logging. Obsidian suggests testing your setup with hostile prompts and malformed tool results. Both sound right. I can't tell you a one person setup can run them cheaply, so I wouldn't build either before the basics above are in place.
Do you need an agent framework?
Label: my rule, with a sourced trade-off. Start with a plain chat window and focused work blocks. Add an orchestration framework only when you can name the specific job it would do that you can't do now. Founderr's ranked list names CrewAI, Lindy and MindStudio, and the trade-off it describes is the one you'd expect: more control means more setup and more ongoing engineering. It's a ranked list, so treat it as opinion.
If you can't finish the sentence "I need a framework to do ___", you don't need one yet.
Your one page to fill in today
- List your apps and the agent (if any) working on each.
- Fill in the access map: reads, changes, keys, needs my yes.
- Write the always-wait list: send, pay, delete, publish.
- Book the weekly 20 minute slot and copy the four questions.
- Add a one line "I ran, nothing to report" habit for every agent.
None of it costs anything. If you tell me which step is blocking you, or the manual task you wish were already automated, I'd like to hear it.
Where to go next
If you'd rather describe your setup or the task you're stuck on, email me and tell me the one thing that keeps breaking.