Null StudioNullStudio

Blog · October 10, 2026 · 9 min read

Software Testing and QA: What to Expect From Your Development Team

By the Null Studio team

TL;DR: Good software testing is not a phase at the end of a project and it is not a percentage on a dashboard. It is a set of habits that let a team change the product quickly without breaking what already works: automated tests around the parts where a mistake is expensive, a test environment that looks like production, a repeatable release process with a way back, and a short acceptance step where your own people confirm the software does the job. As a buyer you do not need to read test code. You need to know where the team has concentrated its testing, how a release reaches users, and what happens when a bug gets through anyway.

Most buyers only think about testing after something goes wrong: a release breaks checkout, or a fix for one bug brings back an old one. By then the question is "why does every change feel risky".

This guide is for the earlier moment, when you are choosing a team and want to know which questions separate a team that tests from a team that says it tests.

What "testing" actually covers

Testing is a family of activities, and each catches a different kind of problem. A healthy project uses most of them, in different amounts.

Automated tests

These are small programs that check the software behaves as expected, run automatically every time someone changes the code.

Unit tests should be many and cheap, end-to-end tests few. A team that covers everything end to end ends up with a slow, flaky suite people learn to ignore.

Manual and exploratory testing

Automated tests only check what someone thought to check. A person using the product with fresh eyes, trying odd inputs, slow connections, the back button, two tabs at once, finds the problems nobody predicted.

User acceptance testing

Acceptance testing is where people from your side confirm that the software does the job in your real workflow. It is not a substitute for the team's own testing. It answers a different question: not "does it work" but "is this what we needed". It goes much faster when the acceptance criteria were written into the brief at the start, which is one reason we recommend doing that in writing a software project brief.

Test where mistakes are expensive

The most common quality mistake is spreading testing evenly. Every product has a few areas where an error costs real money or trust, and many areas where an error is a minor annoyance that can be fixed the same day. Good teams test in proportion to that risk.

Areas that almost always deserve heavy testing:

Internal admin screens, marketing pages and short-lived experiments can usually get lighter testing. That is part of why a prototype costs less than a product that moves money, as we explain in what an MVP costs.

A team that can tell you where its risk is concentrated, and show you that the tests are concentrated there too, is thinking about quality correctly. A team that answers with a coverage percentage alone is measuring something easier.

Some products cannot be tested only in code

Software that touches hardware, devices or other people's environments needs a test plan that goes beyond the code.

When scoping one of these, ask how the team will get access to the real device or environment. The answer affects timeline and budget.

AI features need evaluation, not just tests

Traditional tests assume the same input gives the same output. AI features do not work that way: a model can answer the same question differently twice, and both answers can be acceptable. So AI features are checked with evaluations: a set of realistic inputs, a definition of what a good answer looks like, and a score you track over time, so a prompt change or model upgrade that makes things worse is caught before customers see it.

For voice agents, like the ones behind CallGuard and CallSetter, that means replaying realistic calls (callers who interrupt, change their minds, ask about something the business does not do) and checking that bookings, transfers and handoffs still happen correctly after every change. We cover the production side in AI voice agent monitoring and the product side in adding AI to an existing product.

Environments, releases and a way back

Tests only protect you if the path from code to users is controlled. Four things to expect from any serious team:

  1. A staging environment that resembles production. Same configuration, realistic but non-sensitive data, the same integrations in test mode. "It worked on staging" only means something if staging is honest.
  2. Automated checks on every change. Tests run automatically before code is merged, so broken changes are stopped at the door rather than discovered by users.
  3. Repeatable releases. Anyone on the team can release with the same steps, not one person on one laptop. A release that depends on an individual is a risk we see often in stalled projects.
  4. A rollback plan. When something does break, the previous version can be restored quickly, and database changes are designed so that rolling back does not lose data. For mobile apps, where you cannot instantly pull a release from every phone, that also means staged rollouts and the ability to switch a feature off remotely.

None of this is exotic, and it should be set up in the first weeks, not bolted on before launch.

Bugs will still happen: agree what happens next

No amount of testing produces bug-free software, and a team that promises otherwise is not being straight with you. What you can agree up front:

How this is priced depends on the contract structure, which we compare in fixed price vs time and materials. Whatever the structure, the tests themselves belong to you along with the code, which matters if you ever change teams; see software ownership and handover.

How AI-native development changes testing

Writing tests used to be one of the first things cut when a deadline got tight, because it was slow. Working with AI coding agents changes that trade-off: drafting tests, test data and edge cases is now fast enough that there is little excuse to skip it. That is a large part of how a team can ship quickly without shipping fragile software, which we describe in our ship-in-days playbook.

One caution: a test written by an AI agent still needs a person to judge whether it checks the right thing. A suite that passes while asserting nothing useful is worse than no suite, because it creates false confidence. Deciding what to test is still human work.

Questions to ask a development team

  1. Where do you expect the risk in our product to be, and how will your testing reflect that?
  2. Which tests run automatically on every change, and what happens when one fails?
  3. What does your staging environment look like, and what data will be in it?
  4. How does a release reach our users, and how do you roll back?
  5. For our devices, hardware or integrations, how will you test against the real thing?
  6. If the product includes AI, how will you evaluate it before and after changes?
  7. What do we test during acceptance, and how long should we plan for it?
  8. What happens when we report a critical bug the week after launch?

Good teams answer these in specifics. Vague answers ("we test everything thoroughly") are the warning sign.

The bottom line

Testing is how a team keeps its promise to keep shipping. Focus it where mistakes are expensive, test devices and AI behavior in the conditions they actually run in, control the path to production with a way back, and agree in advance what happens when something slips through. Do that, and every release stops feeling like a gamble.


Want a team that ships fast without breaking what already works? Book a demo and we will walk through where the risk sits in your product and how we would test it. See our work: Fyuel, Raqts, Geonode, LectureNotes AI, ARCortex and more.

FAQ

How much testing does a custom software project need?

It depends on what a mistake would cost. A prototype for a handful of users needs little. Anything that moves money, controls permissions, handles sensitive data or runs inside other companies' software needs heavy testing in those areas. Testing should be built into each piece of work as it is done, not saved for a phase at the end.

What is user acceptance testing and who does it?

User acceptance testing is where people from the client's side confirm the software does the job in their real workflow. It does not replace the development team's own testing. It answers whether the product is what was needed, and it goes fastest when acceptance criteria were written into the project brief at the start.

How are AI features tested?

AI features can give different acceptable answers to the same input, so they are checked with evaluations: a set of realistic inputs, a definition of a good answer and a score tracked over time. That catches a prompt change or model upgrade that makes results worse before customers see it.

Want it built, not just explained?

We design, build and run these systems end-to-end — shipped in days, not months.

Book a demo →

Keep reading