# code_exec — Manual

## What it does
Runs Python in a subprocess and returns stdout (and stderr, prefixed `[stderr]`). The interpreter is **fresh on every call** — no state persists between calls, so import everything you need in each script.

Success means exit code 0. A non-zero exit is returned as `success: false` with the exit code and the full output, which is usually the traceback: read it, it is your own bug, not a tool failure.

## Parameters

| Parameter | Required | Description |
|-----------|----------|-------------|
| `code` | yes | Python source to execute. |
| `working_dir` | no | Directory to run in (see [Limits](#limits)). Defaults to the session's working directory. |
| `timeout` | no | Seconds, default 30, clamped to 1–300. |
| `background` | no | Detach; returns a `task_id` immediately. |

```json
{"code": "import json, pathlib\nprint(json.loads(pathlib.Path('package.json').read_text())['version'])"}
{"code": "print(sum(len(l) for l in open('data.csv')))", "working_dir": "/home/user/project"}
{"code": "import time\nfor i in range(5):\n    print(i); time.sleep(1)", "timeout": 120}
```

## When to use it

Use `code_exec` for anything that is properly a *program*: parsing and transforming data, calculation, text munging, calling a Python library, a quick check of a value. Use `terminal` instead when the task is shell-native — pipes, `grep`/`awk`, system commands, a CLI tool.

A script whose output you then have to read back is often slower than one `filesystem`/`terminal` call. Reach for Python when the alternative is a long chain of tool calls, not by reflex.

## Detaching long runs

Pass `"background": true` for a suite, a large computation or anything that outlives the turn. You get a `task_id` back immediately and the real result arrives as a note at the start of a later turn; check it with the `tasks` tool. A detached call **without** an explicit `timeout` gets 300 s from the executor — so pass `timeout` when the work legitimately needs the ceiling, and remember 300 s is the maximum either way.

## Limits

- Timeout is clamped: `0` becomes 1, `10000` becomes 300, and a non-numeric value falls back to 30. You cannot get more than 300 s in the foreground — use a persistent terminal for genuinely long jobs.
- For a non-admin user the working directory is confined to `user_data/<user_id>/`; an absolute `working_dir` outside it is silently replaced with the sandbox root, and relative paths are resolved inside it. Admin and single-user runs use the client's working directory.
- Temp scripts are written inside that directory and deleted afterwards.
- On timeout the process is killed and you get `Code execution timed out after Ns` with `error: timeout` — a timeout is not a partial result; nothing from the run is returned.

## Common mistakes

- Assuming imports or variables survive between calls. They do not — every call starts clean.
- `print()`-less scripts: nothing is written to stdout means `(no output)`, which looks like a silent failure but is a successful run.
- Using `code_exec` to run `subprocess.run("ls")` when `terminal` or `filesystem list` was the tool.
- Leaving the default 30 s on a test suite that needs 90, then reading the timeout as a broken test.
- Writing to a path outside the sandbox and getting a file that is not where you expected.
