TestinGil - Agile SW Testing Consultancy
Exploring the magic of testing? Trying to perform automation wizardry? Put your Testing Hat on, and I'll help you!
01/10/2026
"You are responsible for our 2.5 hour shutdown in production".
They removed the blindfold.
I was in a dark room. Couldn't see a lot. Then again, I left my glasses on my desk.
"It wasn't me", I said.
"You can't run away from this one. Your name is all over this."
Couldn't run anyway, not in a good shape. And the last coffee was five hours ago.
"Can't be, I've got the tests to prove it".
"You mean those tests?" they asked, pointing at a block of text (with some nice colors, I like dark mode).
"Yup. They pass, Claude wrote the code, the tests, ran them, and promised everything will be ok".
The room went silent. I guess they went for coffee.
"Do you know what the code really does?" They asked.
"Hmm. Sort of". I had to admit, not all the generated code fits inside my head.
"We'll need to take your Claude privileges, until you start reading all this code again."
"Well, I guess we'll meet again here sooner than you think".
They let me go. But I know Claude will sabotage me again. And I'm willing to live with it.
Are you?
https://youtu.be/IAZj1CwKBI0
Flaky Tests From AI-Generated Code: We Hired a Saboteur Ever been handed a flaky test to fix, and the flakiness turned out ...
30/09/2026
Have you ever looked at a legal form, 28 pages long, awaiting your signature in the end? I have.
All you have to do is read everything, find every minute detail that needs fixing, and then you can sign it.
Legal stuff, I'm sure you'll do your best to identify all the small bugs before you sign.
Now imagine having to read things like that ten times a day, but in code font, with as much risk.
And in it there may be some "oopsies", "I apologize" and other confessions. Because our code agents love to tell a story.
New blog - "The Test Was Wrong. Rewriting."
https://testingil.com/2026/09/ai-coding-agent-says-the-test-was-wrong.html
29/09/2026
I love coffee.
I'm one of those weirdos who follow a ritual for making it. (Moka and cold brew, these days).
I have a scale, and I'm using it, and I look at the coffee coming out of the pot, then turn the heat down to control the flow.
Weird, but works. And it works repeatedly, even if there are all kinds of uncontrolled variables. Temperature, beans, grind size, time of day.
And it works because the process has a controller (me). I set the grind size (as much as my grinder lets me), I control the gas flame (as much as my hob lets me) and other stuff. And as long as I pay attention - I get great coffee.
The attention part is crucial here. A controller without attention is like a broken pencil - pointless. (If you know the reference, put a meme in the comments.)
With code agents we abdicate our roles as controllers. We tell ourselves that it's ok, we'll see the end result and tweak then. But then the tests pass, and it looks mostly ok. And we move on to the next task.
I do it too. I gloss over the tests and code, sometimes not even that. We hope it's going to be ok.
With the moka pot, I turn the flame down the moment it spurts. With an agent, I wait until the pot is empty.
In fact we're losing two things - control over the result, and our ability to shorten the feedback loop. If the masses of code are in control, we have no chance of stopping in the middle.
And it does seem like - so what? These are not days or weeks lost. It's hours. Worst case - we'll regenerate.
That's how slop comes out. Like coffee out of an unwatched moka pot.
When did you last stop an agent in the middle, instead of waiting for the end result?
24/09/2026
We've been nagging developers for years to give us tests. Unit tests, API tests - any proof that their code works.
Well, thanks to AI, now we get what we asked for. Unfortunately, the bugs are still in there.
But we can still do something about it.
New blog - link in first comment.
https://testingil.com/2026/09/why-ai-unit-tests-pass-but-code-is-broken.html
23/09/2026
Risks, risks, risks.
Don't you have something good to say about AI?
Sure.
Remember the old days, when you looked at code, and asked: How did it get here? And then you remember reading through docs, and design specs, and then had all this in your mind. Then you added coffee, and got the code.
Now you need to tell a coding agent what to do. It will build it for you. But "what you tell it" can also be recorded.
The intent close to coding, that couldn't be extracted from your mind before, now can be captured, recorded and tracked. In a good way.
See, not all bad. But the rest of the risks are real. So here's the webinar's recording - Link in first comment.
https://www.youtube.com/watch?v=1_BKuS7ZM0g
How to Test AI-Generated Code: The Risks You Inherit [Webinar] Your team ships more code than it used to. Most of it, nobody reads...
17/09/2026
I've had great fun with David Burns on the BrowserStack podcast. We talked about different classical practices in development and testing, like TDD and clean tests.
One of the topics we talked about was self-healing tests. What they are, how they magically heal, and if you can build your own. (Hint: I'm giving a tutorial on how to build your own self-healing test agent at Agile Testing Days 2026).
We talked a lot about how AI fits into different parts of testing, the great benefits and frightful risks. I may have exaggerated with that one.
Watch it here.
https://www.youtube.com/watch?v=D3LNiaWdxw8
Beyond the Green Checkmark: Clean Code, TDD, and AI Testing with TestingGil A green build tells you a lot less than you think. In this episode ...
16/09/2026
I used to trust a green CI pipeline. Now it just gives me trust issues.
Step into my time machine, will you? It's a short ride, I promise.
⠀
Back in the day, as a tester, I didn't naturally trust programmers. I knew the material, well, being one before. But if the devs wrote tests, we became friends. Tests passing means they were doing a good job. Not as well as my evolved dev person. But that's just me.
The point is, I trusted their tests. It saved me time, and I could test more in my very precious testing time.
Back to the future. Now coming pre-installed with AI.
Just last week, we spent three hours debugging a ghost. One of those things that passed through CI, but wreaked havoc on our data.
⠀
Because the genie decided to be "helpful." Everybody, including the coding agent knows that you can't let errors go wandering around unhandled.
The genie agrees.
Instead of letting the system crash (nicely), it added code to just log the error. But by doing this, we now had not-so-integral data in the system.
And of course, it happened more than once. So you can understand the debugging frustration. What can cause all these bits of data? Why didn't the system just stop it (nicely)?
Yes, errors should be handled. But, sometimes it's more important to keep the data uncorrupted.
But the genie missed that in training school.
⠀⠀
We’ve always treated automation as truth. We need that trust in the green to put our efforts somewhere else.
The problem is that tests may be telling the truth. And I'm not sure what this truth means anymore.
⠀
The green isn't as green as yesterday's.
⠀
QA folks and engineering managers: how are you testing the code the AI thinks is "fine"?
15/09/2026
Who's afraid of the big bad AI?
Turns out a lot of you.
My webinar tomorrow "Testing Code Nobody Reads: Handling the Risks of AI-Generated Code" covers some of the things people mentioned - "Too much unreviewed code", "Unwanted AI additions" and more.
This is what most people see. And I have news for you - it gets scarier. I've got a lot more risks to talk about.
Good news too - we can do something about it.
Join me tomorrow. We're stronger in big groups.
Sep-16th, 3PM CEST / 9AM EDT.
https://us02web.zoom.us/webinar/register/8517677800361/WN_2G3Dv53sSeaDaPjtwOXeRQ
10/09/2026
I used to review code. A lot. Still do.
I didn't always understand everything. But I had a fallback.
Whether it was developer code, or automated tests. If I didn't understand something, or why a certain choice was made, or even the order of things - I always could go to the original writer, and ask - what did you mean by that?
Why have you selected this option? What's behind this design?
What do you expect to see when you debug it?
These questions help me understand the code better, and approve the PRs.
But these days, I don't have the luxury. AI writes everything, and I don't have anyone to ask. The genie has made its choice, and if I ask "why didn't you go with X?" the answer is usually "Good catch! Let me change it for you".
Most of our knowledge of our system sits in people's heads. It now moves at the speed of AI into nothingness. Who will be able to explain these things in two years' time?
We can do something about it. I'll talk about it in my next webinar "Testing Code Nobody Reads: Handling the Risks of AI-Generated Code" next week.
September 16th · 3PM CEST / 9AM EDT · 60 minutes. Free.
Register here:
https://us02web.zoom.us/webinar/register/8517677800361/WN_2G3Dv53sSeaDaPjtwOXeRQ
09/09/2026
Why do we do code reviews?
Lots of answers to that. But how about knowing that if your teammate went to get married in Tahiti, and didn't invite you, and a bug came in - you'll be prepared to take on that bug.
In code reviews, we absorb the code changes, and understand them to be ready for the next bug or next feature.
I thought I was doing that with my code reviews for my code agent. Turns out it's not so simple.
New blog:
https://testingil.com/2026/09/code-review-for-ai-generated-pull-requests.html
Click here to claim your Sponsored Listing.
Category
Contact the business
Website
Address
Hod Hasharon
4529546
Alerts
Be the first to know and let us send you an email when TestinGil - Agile SW Testing Consultancy posts news and promotions. Your email address will not be used for any other purpose, and you can unsubscribe at any time.