TL;DR: Good software testing is not a phase at the end of a project and it is not a percentage on a dashboard. It is a set of habits that let a team change the product quickly without breaking what already works: automated tests around the parts where a mistake is expensive, a test environment that looks like production, a repeatable release process with a way back, and a short acceptance step where your own people confirm the software does the job. As a buyer you do not need to read test code. You need to know where the team has concentrated its testing, how a release reaches users, and what happens when a bug gets through anyway.
Most buyers only think about testing after something goes wrong: a release breaks checkout, or a fix for one bug brings back an old one. By then the question is "why does every change feel risky".
This guide is for the earlier moment, when you are choosing a team and want to know which questions separate a team that tests from a team that says it tests.
What "testing" actually covers
Testing is a family of activities, and each catches a different kind of problem. A healthy project uses most of them, in different amounts.
Automated tests
These are small programs that check the software behaves as expected, run automatically every time someone changes the code.
- Unit tests check one piece of logic in isolation: a price calculation, a date rule, a permission check. They are fast and cheap, and they are where business rules belong.
- Integration tests check that pieces work together: the app and its database, your system and a payment provider, a booking flow and a calendar.
- End-to-end tests drive the product the way a user would, through the real screens, for the handful of journeys that must never break, such as sign-up, purchase or submitting a form.
Unit tests should be many and cheap, end-to-end tests few. A team that covers everything end to end ends up with a slow, flaky suite people learn to ignore.
Manual and exploratory testing
Automated tests only check what someone thought to check. A person using the product with fresh eyes, trying odd inputs, slow connections, the back button, two tabs at once, finds the problems nobody predicted.
User acceptance testing
Acceptance testing is where people from your side confirm that the software does the job in your real workflow. It is not a substitute for the team's own testing. It answers a different question: not "does it work" but "is this what we needed". It goes much faster when the acceptance criteria were written into the brief at the start, which is one reason we recommend doing that in writing a software project brief.
Test where mistakes are expensive
The most common quality mistake is spreading testing evenly. Every product has a few areas where an error costs real money or trust, and many areas where an error is a minor annoyance that can be fixed the same day. Good teams test in proportion to that risk.
Areas that almost always deserve heavy testing:
- Money. Anything that calculates, charges, refunds or reconciles. In an accounting platform like Fyuel, which brings customers, suppliers, tankers, ledgers and banks into one system, the ledger logic is where we would want the densest tests, because a rounding or posting error silently corrupts every report built on top of it. We go deeper in building financial software.
- Permissions. Who can see and change what. A bug here is a data leak, not a cosmetic issue.
- Integrations. Every connection to an outside system is a place where the other side can change, slow down or return something unexpected. Tests should cover the failure cases, not just the happy path. See software integration projects.
- Data changes. Migrations and imports that touch existing records, because a mistake is hard to undo once it has run against production.
Internal admin screens, marketing pages and short-lived experiments can usually get lighter testing. That is part of why a prototype costs less than a product that moves money, as we explain in what an MVP costs.
A team that can tell you where its risk is concentrated, and show you that the tests are concentrated there too, is thinking about quality correctly. A team that answers with a coverage percentage alone is measuring something easier.
Some products cannot be tested only in code
Software that touches hardware, devices or other people's environments needs a test plan that goes beyond the code.
- Hardware and sensors. Raqts, the racquet-sport wall that responds to a player's hits, depends on computer vision and connected hardware. Logic can be tested in code, but detection has to be checked against recorded and live sessions in real lighting with real players, because that is where it fails. We cover this in computer vision product development.
- Software that ships inside other apps. The cross-platform SDKs we built for Geonode run inside apps we do not control, on many operating systems and versions. Testing means a matrix of platforms, plus checks that a new version does not break customers still running the old one. More in cross-platform SDK development.
- Headsets and XR. VR and AR builds such as Nystag, which runs eye-tracking diagnostics on the Vive Focus 3, have to be tested on the actual device for frame rate, tracking and comfort, and retested whenever the headset's software updates, as we explain in XR app maintenance.
When scoping one of these, ask how the team will get access to the real device or environment. The answer affects timeline and budget.
AI features need evaluation, not just tests
Traditional tests assume the same input gives the same output. AI features do not work that way: a model can answer the same question differently twice, and both answers can be acceptable. So AI features are checked with evaluations: a set of realistic inputs, a definition of what a good answer looks like, and a score you track over time, so a prompt change or model upgrade that makes things worse is caught before customers see it.
For voice agents, like the ones behind CallGuard and CallSetter, that means replaying realistic calls (callers who interrupt, change their minds, ask about something the business does not do) and checking that bookings, transfers and handoffs still happen correctly after every change. We cover the production side in AI voice agent monitoring and the product side in adding AI to an existing product.
Environments, releases and a way back
Tests only protect you if the path from code to users is controlled. Four things to expect from any serious team:
- A staging environment that resembles production. Same configuration, realistic but non-sensitive data, the same integrations in test mode. "It worked on staging" only means something if staging is honest.
- Automated checks on every change. Tests run automatically before code is merged, so broken changes are stopped at the door rather than discovered by users.
- Repeatable releases. Anyone on the team can release with the same steps, not one person on one laptop. A release that depends on an individual is a risk we see often in stalled projects.
- A rollback plan. When something does break, the previous version can be restored quickly, and database changes are designed so that rolling back does not lose data. For mobile apps, where you cannot instantly pull a release from every phone, that also means staged rollouts and the ability to switch a feature off remotely.
None of this is exotic, and it should be set up in the first weeks, not bolted on before launch.
Bugs will still happen: agree what happens next
No amount of testing produces bug-free software, and a team that promises otherwise is not being straight with you. What you can agree up front:
- Severity levels. What counts as critical (the product is down, money or data is wrong), what counts as serious and what can wait for the next release.
- Response times for each level, and who you contact.
- What is covered. Many teams include a warranty period after launch for defects in delivered work. Agree what that period covers and what counts as new work instead.
- A regression test for every serious bug. When a bug is fixed, a test should be added so the same bug cannot quietly return.
How this is priced depends on the contract structure, which we compare in fixed price vs time and materials. Whatever the structure, the tests themselves belong to you along with the code, which matters if you ever change teams; see software ownership and handover.
How AI-native development changes testing
Writing tests used to be one of the first things cut when a deadline got tight, because it was slow. Working with AI coding agents changes that trade-off: drafting tests, test data and edge cases is now fast enough that there is little excuse to skip it. That is a large part of how a team can ship quickly without shipping fragile software, which we describe in our ship-in-days playbook.
One caution: a test written by an AI agent still needs a person to judge whether it checks the right thing. A suite that passes while asserting nothing useful is worse than no suite, because it creates false confidence. Deciding what to test is still human work.
Questions to ask a development team
- Where do you expect the risk in our product to be, and how will your testing reflect that?
- Which tests run automatically on every change, and what happens when one fails?
- What does your staging environment look like, and what data will be in it?
- How does a release reach our users, and how do you roll back?
- For our devices, hardware or integrations, how will you test against the real thing?
- If the product includes AI, how will you evaluate it before and after changes?
- What do we test during acceptance, and how long should we plan for it?
- What happens when we report a critical bug the week after launch?
Good teams answer these in specifics. Vague answers ("we test everything thoroughly") are the warning sign.
The bottom line
Testing is how a team keeps its promise to keep shipping. Focus it where mistakes are expensive, test devices and AI behavior in the conditions they actually run in, control the path to production with a way back, and agree in advance what happens when something slips through. Do that, and every release stops feeling like a gamble.
Want a team that ships fast without breaking what already works? Book a demo and we will walk through where the risk sits in your product and how we would test it. See our work: Fyuel, Raqts, Geonode, LectureNotes AI, ARCortex and more.