Well met!
This week, I listened in on a whiteboarding session with Aulendur's engineering team as they discussed their approach to LLM-written tests.
Time vs. Tests
Time means everything to a small team. The accelerated development capabilities of modern LLM coding companions are crucial to the success of creating software solutions at today's expected pace. Does that mean that in the interest of speed, that we should wholesale trust LLM-written code and LLM-written unit tests? The answer the team landed on is: no, absolutely not.
But how can we leverage the acceleration of LLM-written unit tests and trust that new generated code will improve our codebase, not worsen it? Should we even be spending time writing unit tests?
We have previously determined that all critical* code needs to be evaluated by a human/engineer. Essentially, someone in the organization needs to comprehend and understand what the code does. This is ultimately critical for the sustainment of our projects. As codebases grow over time they also grow in complexity, scope, and size. Since LLMs have been shown to fail with scale, there is a meaningful risk that an LLMs capacity to perform “good” engineering decreases over the life of a project. If the engineers have offloaded comprehension of the project to the LLM, they also accept the risk that they may find themselves responsible for sustaining a project they do not understand, with no help from the LLM. Until this risk can be mitigated with certainty, we will only allow critical* applications to run on software that someone in our organization understands and can defend - though they may use AI tools to assist them in its generation.
Unit tests present a unique problem however. Writing good tests has long been a clear sign of a high performance team. A well written test afterall, requires a certain depth of thought about the goal of a function/module/etc and the edge cases for potential problems. When an LLM writes unit tests, even if they are very good, it does not represent the same progress as when a human has written it. The value of good tests therefore needs to be split:
- There is the direct value to the codebase, which is that we now have an automated method to reduce possible errors.
- There is also value that is in the engineer who writes it.
As a community, when we see a good test, we assign a positive value to that which can no longer be taken for granted. It has been a useful heuristic to combine these two values, but as with many such issues in this new LLM hellworld™ we need to unpack our feelings and reconsider how we assign value.
So, do we really need to understand all the tests that the LLM writes in the same way that we enforce understanding in the “regular” code?
Free Cheese
If we’re being honest, we didn’t always have the time or inclination to write ALL the tests we would have liked to. However, we did still write good tests. Let’s not lose that. Whatever tests we WOULD have written in the "old world" are tests we should still handwrite. Let’s make sure we get the goodness out of the process and artifacts of writing those awesome end-to-end high level tests. But any tests that are “free cheese” that the LLM writes that we wouldn’t have written? Why should we throw them away? Usually, these will help us catch later changes to the code that might break something. Worst case, the LLM-written test catches a false positive, we take a closer look at the test at that time and determine if it’s a poorly written test and either fix it, or delete it.
Let’s not throw the molten steel out with the slag.
Big takeaways:
- Code tests are still important, whether we write them or an LLM writes them, they should still be made
- We should give LLM-written tests a hearty once over
- We should still hand-write tests we would have written ourselves before LLMs, like core functionality or for flagging dangerous scenarios
- Engineers should still manually test their app to validate end to end functionality when possible
- End to end tests when possible are still the most valuable tests