Run Bob and Walk Away
What it takes to run an AI agent in a container on a schedule
Most developers use IBM Bob interactively. You open a session, ask something, close the tab. That's fine, but it means Bob sits idle for the sixteen hours a day you're not at your keyboard.
I wanted more. I wanted to schedule a job, go home, and wake up to a Slack summary of what Bob had done overnight. Here's what I built, what broke, and what I learned.
Install Bob Like a Dependency
The first thing I had to figure out was how to get Bob into a container. It's easier than it sounds: Bob Shell installs from a tarball URL straight in package.json:
"bobshell": "https://s3.us-south.cloud-object-storage.appdomain.cloud/bob-shell/bobshell-2.0.5.tgz"
This is where the published install script for Bobshell pulls from, and we can save time by directly adding that to our package.json file. If your container can run Node, it can run Bob.
From there, you need a way to call Bob from a script. Bob Shell's CLI accepts a prompt over stdin, so the wrapper is a thin one:
const bob = spawn("node", ["node_modules/.bin/bob", "run", "--trust", "--max-cost", "50"], {
stdio: ["pipe", "inherit", "inherit"],
});
createReadStream("prompts/my-task.md").pipe(bob.stdin);
We wrapped this in an askBob(promptFile) utility so the orchestrator reads cleanly. Bob runs the prompt, streams its output to the terminal, and the promise resolves when the session ends. Bob is just a process you spawn and pipe a file into.
The Naive Version Breaks Immediately
Once I had Bob running in a container, I scheduled it as an OpenShift CronJob and called it done. That lasted about one run.
The hang. Bob ran a test suite but Jest didn't exit cleanly. The execute_command call had no timeout, so the pod just sat there... for two hours and forty minutes, until OpenShift's deadline killed it. I had no error in the logs, no Slack message in the morning, and a mystery to solve.
The silent crash. A later run threw an exception halfway through. The job exited. Again, no helpful Slack message. I found out when I noticed no PRs had appeared overnight.
The bad output nobody caught. Bob wrote a structured JSON file that a later phase depended on but the JSON had a missing field. The next call ingested it quietly and produced nonsense downstream. I didn't notice for two runs.
None of these feel like problems when you're watching Bob work interactively. You catch them. Unattended, each one is a wasted run.
Three Things That Change When Nobody's Watching
Debugging those failures made me realise running in a container is harder than it looks. I need to recreate all the subtle error checking I do throughout the day, but in an automated way. Three things are fundamentally different.
You can't intervene
When Bob is running live, you are the guardrail. If a command hangs, you can quickly kill it and move on. The system has to do that when running unattended.
I fixed the hang with a PreToolUse hook. The hook is a small script that intercepts every execute_command call before it runs and blocks any that don't include a timeout_seconds value:
if (payload.tool_name !== "execute_command") process.exit(0);
if (input.background === true) process.exit(0);
if (!input.timeout_seconds) {
process.stderr.write(
"execute_command blocked: timeout_seconds is required. " +
"Re-issue the call with an explicit timeout_seconds value."
);
process.exit(2); // blocks the call
}
Bob sees the denial message, re-issues the call with an explicit timeout, and continues. There is a slight penalty for Bob, costing time and tokens for some commands that honestly don't need a timeout, but the benefit is worth it.
The key idea here is that hooks let you encode constraints as enforced policy. The timeout requirement fires on every call, forever, regardless of what prompt Bob is following.
You won't see failures
Invisible failures are incredibly frustrating. The fix is putting another Bob prompt in a finally block so it runs regardless of what happens:
try {
await main();
} catch (err) {
console.error(`Fatal: ${err.message}`);
} finally {
await askBob("prompts/debug-to-slack.md"); // always runs
}
Right now, I'm using a Slack summary here. Win or fail, you wake up to a message. In the future, we can expand to self-recovery: the same block that posts a summary can ask Bob to diagnose what went wrong and attempt a fix before escalating.
Context accumulates in ways you don't notice
Running interactively, you naturally start a new session for each unrelated task. In a long scheduled job, that doesn't happen unless you design for it. By the time Bob is working on repository ten, it's carrying context from repositories one through nine: assumptions, partial state, details that no longer apply.
Phase-based orchestration solves this. Structure the job as a sequence of askBob() calls, each with a focused prompt:
await askBob("prompts/01-assess.md");
await askBob("prompts/02-fix.md");
await askBob("prompts/03-review-ci.md");
await askBob("prompts/04-summarise.md");
Each call is a fresh context. Bob finishes one task, the window closes, the next starts clean. Individual phase failures can be caught and skipped without aborting the whole run. And when something goes wrong, the logs tell you exactly which phase failed.
Teach Bob Your World with Skills
Bob out of the box knows how agents work. He doesn't know how your CI pipeline works, what your credential rotation looks like, or what your test coverage gate requires.
Skills close that gap. They're markdown files you add to .bob/skills/ where Bob can read them before starting a task. A skill for vulnerability remediation tells Bob which tools to call and in what order. A skill for reviewing CI results tells Bob which failures are transient. A skill for unit test generation tells Bob what "done" looks like in your codebase.
The difference between Bob with skills and Bob without is the difference between a contractor who knows your codebase and one who's starting fresh every time. Without skills, Bob guesses at your conventions. With skills, Bob executes them.
Skills compound too. Every one you write makes every future job better. Write them to encode the multi-step processes with non-obvious details, the kind of thing you'd explain to a new team member on their first day.
What I'd Take From This
The specific prompts and jobs are ours. They'll change as our needs change. But the architecture underneath is reusable:
- Install Bob as a dependency, version-pinned in source control
- Use hooks to enforce constraints Bob can't enforce itself
- Put the final summary in a
finallyblock - Structure work as phases so failures are isolated and logged
- Teach Bob your conventions with skills before the first run
The next post covers what we actually built on top of this: running Bob every night to work through a vulnerability backlog across a fleet of repositories. That's where the payoff is.