If you use an AI helper every day for the same routine, checking messages, reviewing what happened overnight, writing up a report, the cost creeps up quietly. The natural instinct is to trim whatever the internet says is the usual culprit. That is exactly what we almost did at Balay ni Bruno & Co.
A lot of public guidance says the big cost is all the tools you plug into the AI. Some sources talk about huge overheads from those connections on every single turn. We had more than 20 of them connected, so it seemed obvious. Then we measured a real session that had been running for about eight shifts, and the answer was somewhere else entirely.
What the Measurement Showed
Where the AI's reading went in one real session that had run about eight shifts.
The tool connections were a real problem, just not a cost problem. Each session used about 1.1 GB of the computer's memory because of them. The money problem was different: one long running session was carrying every shift's raw data, and every new turn sent all of it to the AI again.
Measure before you cut. The thing everyone warns you about may be a small slice of your bill, and the real cost may be sitting in plain sight in your own chat history.
The Five Changes That Work, in Order
Start a new session for every run. Keep what carries over in written notes, not in the chat.
A small script shrinks raw data before the AI ever sees it.
Yes or no questions are answered by a script, not by the AI.
Simple sorting and summaries go to the cheapest model that does them well.
1. A fresh session for every run
This one is free. Instead of one session that lives for days, every run starts clean. Anything that needs to carry from one day to the next lives in written files the AI can read when it needs them, not in a scroll of old messages it rereads on every turn.
2. Never paste raw data into the AI
A small script digests it first. Ours took a busy day of 230 chat messages from 38,434 characters down to 18,162, a cut of 53%. It did that by dropping greetings, merging several messages in a row from the same person, shortening links, and trimming our own updates harder than the client's words.
Pasting the raw chat
- Every hello and thank you included
- Five short messages in a row, kept as five
- Long links in full
- Our own long updates repeated back
- 38,434 characters for one day
A digested summary
- Greetings dropped
- Runs from one person merged into one
- Links shortened
- Client words kept exactly, our echoes trimmed
- 18,162 characters for the same day
The rule of thumb: the client's exact words matter. Your own repeated updates do not. Trim yours first.
3. Let code answer the yes or no questions
Does a file exist? Is a scheduled job still switched on? What is the balance on an account? Does an old key still work? None of these need an AI to think. A script checks each one and prints DONE, NOT DONE or UNKNOWN, with the evidence. The AI then only writes about what is new. The very first time we ran our checking script, it caught two daily jobs that had stopped, which a manual check had missed the same night.
4. Use the cheapest capable model for simple passes
Summarising and sorting do not need the most powerful model. Save the top model for the parts that need judgement.
5. Keep the start of every request the same
The AI was already reusing the unchanged opening part of each request within a session, which is cheaper. The saving comes from keeping that opening stable, not from changing any settings.
What We Tried and Set Aside
We also looked at two other options on the same day and decided against both. Sending the work in a slow overnight batch takes up to 24 hours to come back, and this routine needs answers while we work. Running an AI model on our own computer would have meant buying hardware and waiting longer for every answer, for a job that runs twice a day.
How We Keep the Checks Safe
Read only, by design. The checking scripts can look but cannot change, delete or send anything. That is what makes them safe to run on their own, over and over, even while other systems are busy.
- Unknown is never a pass. If a check fails to read something, it reports that it could not look. It never reports that everything is fine.
- It works from anywhere. A scheduled run does not start where you were working, so the scripts cannot depend on that.
- Re-measure after a full shift. Check the real usage afterwards rather than trusting an estimate.
Common Questions
Is it the extra tool connections that make a daily AI routine expensive?
Not in our case. A lot of public advice says the tools you connect to the AI are the big cost. When we measured a real session that had run for about eight shifts, all of our 20 plus tool connections together were only 3% of what the AI was reading. The conversation itself, meaning pasted chats, long command output and long reports, was 83%. The connections did use a lot of computer memory, but that is a different problem.
What is the single cheapest fix?
Start a fresh session for every run. It costs nothing. One long running session carries every earlier day's data and sends it all again on every turn. Keep what needs to carry over in written notes and files instead of in the chat history.
Do I need the most powerful AI model for everything?
No. Simple jobs like summarising and sorting do not need the top model. Send those to the cheapest model that does them well, and save the strongest model for the parts that need real judgement.
Is it safe to let these cost saving checks run on their own?
Ours are built to only read, never change, delete or send anything, so they are safe to run unattended and as often as needed. They also never treat a failed check as a pass: if a check could not look, it says so plainly instead of reporting that everything is fine.
Key Takeaways
- Measure before you cut. In our real session the conversation was 83% of the reading, and all 20+ tool connections were 3%.
- Start a fresh session for every run and keep continuity in written notes. It is free.
- Digest raw data with a script first. Our daily chat summary shrank by 53%.
- Let code answer yes or no questions, and let the AI write only about what is new.
- Send simple passes to the cheapest model that does them well.
- Build checks that only read, and never let "could not look" count as "fine".