Guide sources and teaching inputs checked . Vendor directory claims retain their own review dates.
Check what the agent can actually reach
Inspect filesystem grants, attached files, connected drives, search tools, shell access and network access. An agent may read beyond its starting directory or send code and file contents to a cloud service. “Local editor” and “empty folder” do not establish that processing stays local.
Use a dedicated workspace with fictional inputs. Remove unnecessary connectors and restrict access through real controls, not only a prompt asking the assistant to avoid sensitive files.
Keep the first build fictional
Try Codex with a concrete brief: “Build a visit-preparation page from this fictional CSV. Show missing values, keep the source file unchanged and include an example I can check by hand.” Ask it to show the result, explain its assumptions and test an incomplete record before adding any integration.
Write a short brief, name what the tool may read and change, and have it create a new output instead of overwriting source files. Open the output and compare it with the teaching inputs. Check unexpected reads, external calls, errors and incomplete output.
Before real clinical use, separately review the code, execution environment, agreements, access and data handling. Running generated code yourself does not automatically make it safe for PHI.
Separate Codex Local from Codex cloud
OpenAI’s guide includes ChatGPT for Healthcare, ChatGPT for Clinicians and Regulated workspaces with an applicable BAA and required Codex access. Confirm the actual account and enablement before PHI use. API-key harnesses instead require BAA-covered API Services with Modified Retention and the required provisioning. Codex cloud is excluded.
A harness does not inherit permission to send PHI to every tool it can call. Review local files, logs, repositories, browser activity and each MCP server, plugin or connected service. Start with fictional inputs and the least access needed, then assess the full data path before real use.
Copy a prompt to try
Brief for a small internal tool
Build & automate
Build a small tool that does this: {{what_it_does}}
Before writing any code:
- restate what you think I asked for, in one paragraph
- list the decisions you are about to make on my behalf
- ask me about anything you would otherwise guess
Then build it to run on my own machine, with no database and no accounts.
Use made-up sample data. Show me how to run it and how to stop it.
Build with fictional rows in a dedicated workspace. Inspect actual file, network and connector permissions: an empty starting folder does not confine an agent. Separately approve the data path and code before real clinical use.
I do this by hand every week: {{task}}
Ask me what the input looks like and what the output has to look like.
Then write a script that does it, and make it:
- refuse to run if the input is not what it expects
- write to a new file, never over the original
- print exactly what it changed
Finish by telling me what to check the first few times before I trust it.
Takes a description of the work, not the files themselves. Run it on copies until you trust it. If the real files hold patient information, then the machine you run it on and the tool you built it with both need to be ones allowed to see that.