Hi everyone. If you work in QA, over the last few months you have probably heard the word "agentic" more times than you can count. It shows up in talks, job postings, tool launches and every LinkedIn post about the future of testing. And as with every buzzword, it is often used without explaining what it actually means.
In this article I want to calmly sort out the topic: what agentic QA is, how it differs from the automation we already know and from chat-style AI assistants, how an agent works under the hood, what tasks it can really do today, what its risks are and, above all, what changes for us as QAs. At the end I leave you a guide to start trying it without losing your mind.
Spoiler: no, AI is not coming to replace QA. But it is changing the way we work a lot, and it is worth understanding before it runs us over.
What is agentic QA?
Agentic QA (also called agentic QA or agentic testing) is a way of doing quality assurance in which artificial intelligence agents perform testing tasks autonomously or semi-autonomously: they understand a goal, decide which steps to follow, use tools to carry them out, observe the result and adjust their plan on the fly.
The key word is agent. An AI agent is not simply a language model that answers questions. It is a system that combines several pieces:
- A language model (LLM) that reasons, interprets instructions and decides what to do.
- Tools that let it act on the world: open a browser, click, send a request to an API, read a file, run a command, query a database or create a ticket in Jira.
- A work loop in which the agent thinks, acts, observes what happened and thinks again, until it achieves the goal or realizes it cannot.
- Memory and context: information about the project, the requirements, the history of previous runs and the team's rules.
- A goal, which instead of being a fixed sequence of steps is something like "verify that a new user can sign up and complete their first purchase".
That last difference is the most important one. In traditional automation we tell the machine how to do each thing: "find the element with this selector, type this text, click here, check that this message appears". In the agentic approach we tell it what we want to achieve, and the agent figures out the how, adapting if the screen changed, if an unexpected popup appeared or if the button now has a different name.
To make it clear from the start: agentic QA does not mean "QA without humans". In practice, the best results come when the agent does the repetitive, heavy work and the QA defines the strategy, reviews what the agent produces, makes decisions and brings the judgment the machine still lacks. This is what the industry calls human-in-the-loop: the human stays in the loop, but in a different role.
From scripts to agents: a bit of history
To understand why agentic QA is a big shift, it helps to look at how testing has evolved over the last few decades. Simplifying quite a bit, we can think of five stages.

1. Manual testing
It all started with people running test cases by hand, following a document or a spreadsheet. Manual testing is still essential, especially for exploratory testing, usability and anything that requires human judgment. But it scales poorly: testing the same thing on every release is slow, expensive and boring, and boredom is the enemy of attention.
2. Scripted automation
Then came tools like Selenium, and later Cypress and Playwright, that let us write code that runs the tests. This revolutionized regression testing: what used to take days now runs in minutes inside a CI/CD pipeline. The problem is that scripts are brittle. If a selector, a text or the order of a screen changes, the test breaks even though the functionality still works fine. Anyone who has maintained a large suite knows that a big part of the time goes into fixing tests, not finding bugs.
3. Frameworks, BDD and low-code
To make automation more maintainable and accessible, patterns such as Page Object appeared, approaches such as BDD with Cucumber and Gherkin, and low-code tools that record user actions. They helped enormously to organize the work and bring automation closer to less technical profiles, but the underlying logic stayed the same: steps defined in advance, executed to the letter.
4. AI as an assistant
With the arrival of modern language models, we started using AI as a copilot: we ask a chat to generate test cases from a user story, write us a script, explain an error or put together test data. It is a huge leap in productivity, but the human still acts as the intermediary for everything: copying, pasting, running, reviewing, asking again.
5. AI as an agent
The current stage is the agent. Instead of giving us text for us to use, the agent acts directly: it reads the user story, generates the test cases, opens the browser, runs the tests, gathers the evidence, analyzes the failures and leaves us a report. We move from executing to supervising, from writing every step to defining goals and quality criteria.
No stage completely replaced the previous one. Today manual testing, scripts, BDD, assistants and agents all coexist. The key is knowing which tool to use for each problem.
AI assistant vs. AI agent
This is a very common confusion, so it is worth pausing on. Not everything with AI is agentic. Many tools sold as "agents" are really assistants with good marketing. This table summarizes the main differences:
| Aspect | AI assistant | AI agent |
|---|---|---|
| How it works | Answers each request and waits for the next one. | Receives a goal and chains several steps on its own. |
| Action | Produces text, code or suggestions. | Uses tools: browser, terminal, APIs, files. |
| Adaptation | Depends on you telling it what happened. | Observes the result of its actions and corrects course. |
| QA role | Operator: asks, copies, runs. | Supervisor: defines, reviews, approves. |
| Example | "Write me test cases for this login." | "Test the login, document what you find and report the bugs." |
A simple way to tell them apart: if the tool needs you to act as the bridge between what it tells you and what happens in the system, it is an assistant. If it can interact with the system by itself, observe results and decide the next step, it is an agent.
Mind you, an assistant is not "worse". For many tasks, such as thinking up scenarios, reviewing the wording of a test case or understanding a concept, an assistant is perfect and cheaper. The agent shines when the task has many steps, requires interacting with real systems and is repeated often.
How a QA agent works under the hood
You don't need to be a machine learning expert to understand how an agent works. Its behavior can be explained with a fairly intuitive loop, very similar to how a human tester works.
The loop: think, act, observe
- Understand the goal. The agent receives an instruction ("validate the password recovery flow") together with context: the user story, the acceptance criteria, the URL of the test environment, the test credentials.
- Plan. The model reasons about the steps it needs: go to the login screen, find the "forgot my password" link, enter a valid email, check the message, check the test mailbox, follow the link, set a new password and try to log in.
- Act. It uses a tool to perform the first step, for example opening the browser at the given URL.
- Observe. It receives the result of the action: the page content, a screenshot, the API response or an error message.
- Adjust. With that information it decides the next step. If a cookie banner appeared, it closes it. If the link moved, it looks for it. If it found something odd, it records it.
- Repeat until the goal is achieved, it runs out of options or it reaches a defined limit (of steps, time or cost).
- Report. It puts together a summary of what it did, what it found and the corresponding evidence.
This pattern of reasoning, acting and observing in a loop is the basis of almost every agent today. What makes it powerful is that the agent does not need every step written in advance: it works things out on the fly, as a person would.

Tools and the MCP protocol
A language model on its own cannot open a browser or query a database. For that it needs connected tools. For a long time every platform solved this in its own way, until the Model Context Protocol (MCP)appeared: an open standard that defines how an agent connects to external tools and data sources.
For the QA world, MCP is great news: today there are MCP servers to control browsers with Playwright, work with repositories, query databases or manage tickets. That means a single agent can, in one session, read the story in your management tool, run the test in the browser, query the database to check the data was saved correctly and leave the bug documented. All without you having to copy and paste between windows.
Context: the agent's fuel
The quality of an agent's work depends heavily on the context we give it. An agent that does not know the business rules will test obvious things and miss what matters. That is why there is more and more talk of context engineering: preparing clear documentation, well-written acceptance criteria, domain glossaries, team conventions and examples of good test cases and good reports.
Notice something interesting: everything the agent needs to do its job well is exactly what QA has always advocated. Clear requirements, testable acceptance criteria, up-to-date documentation. The agentic approach does not replace good practices; it makes them even more necessary.
One agent or several
Many current systems do not use a single agent that does everything, but a team of specialized agents coordinated by an orchestrator. For example: one agent analyzes stories, another designs test cases, another generates data, another runs tests in the browser and another builds reports. Each one has instructions and tools limited to its task, which usually gives better results than a generalist agent and makes it easier to understand what failed when something goes wrong.
It is the same approach we use in the k0lmenIA agents: eleven agents with well-defined responsibilities, from story analysis to the closing report, designed so that a manual QA can use them by chatting in Spanish.
What a QA agent can do today
Let's separate the hype from reality. These are tasks agents already do reasonably well today, always with human supervision:
Analysis of requirements and user stories
An agent can read a story and detect ambiguities, missing acceptance criteria, contradictions or edge cases nobody considered. It is a very concrete way to apply Shift-Left: finding problems before a single line of code exists. The questions the agent generates are excellent input for refinement meetings.
Test case design
From the requirements, the agent can propose test cases applying classic techniques such as equivalence partitioning, boundary values or decision tables, and write them in your team's format: a spreadsheet, a test case manager or Gherkin scenarios. The result is not perfect, but as a first draft it saves a huge amount of time and usually covers scenarios you would miss out of fatigue.
Test data generation
Creating realistic, varied and mutually consistent data is a tedious task that agents handle very well: users with different profiles, valid and invalid addresses, test cards, unusual character combinations, files with edge-case formats. They can also generate data that respects the business rules, something that is hard with traditional random generators.
Running E2E tests in the browser
Connected to a browser, an agent can run an end-to-end flow, interpret the screen, fill in forms and verify results, without depending on hand-written selectors. This is especially useful for smoke tests, quick checks after a deploy and guided exploratory testing.
API testing
From an OpenAPI contract or a Postman collection, an agent can propose test cases, send the requests, validate status codes, response structures and business rules, and document what it finds. If you want to review the fundamentals, the blog has guides on Postman and Swagger, and an HTTP status code reference that will come in handy when reviewing what the agent reports.
Self-healing of automated tests
When an automated test fails because a selector or a text changed, an agent can analyze the error, inspect the page, propose the fix and, if the team allows it, apply it. This directly tackles the biggest pain of traditional automation: maintenance. That said, you have to be very careful that the agent does not "fix" a test that failed because of a real bug.
Failure analysis and triage
When a suite of hundreds of tests fails in the pipeline, someone has to look at the logs and decide what is a bug, what is an environment problem and what is a flaky test. Agents can group similar failures, identify the likely cause and prioritize what to review first. For teams with large suites, this saves hours every week.
Bug and results reports
From the messy notes of a test session, an agent can write clear bug reports, with steps to reproduce, expected result, actual result, suggested severity and evidence. It can also consolidate the results of a test round into an executive report with metrics and a go-live recommendation.
Other growing tasks
- Pull request review from a testing perspective: what changed, which tests are missing, what risks it introduces.
- Accessibility: detecting contrast problems, missing labels and keyboard navigation issues.
- Visual testing: comparing screens to find the differences that matter and ignore the ones that don't.
- Test selection: deciding which subset of the regression suite to run based on the changes in each commit.
A working day with agents
To make all this concrete, let's imagine Sofi, a QA on a team building a booking app. This week a new feature enters the sprint: cancelling a booking with a partial refund depending on how far in advance it is cancelled.

In the morning, Sofi gives the user story to the analysis agent. Within minutes she gets a list of questions: what happens if the cancellation occurs exactly at the 48-hour limit? Is the refund calculated on the price with or without taxes? What does the user see if the payment method has already expired? Sofi discards two questions that don't apply, adds one of her own and takes the rest to the refinement meeting. The Product Owner adjusts the acceptance criteria.
Before lunch, she asks the design agent to generate test cases from the updated criteria. She gets forty cases. She reviews them, removes duplicates, fixes two that misinterpreted a rule and prioritizes. She asks the data agent to prepare bookings with different lead times and payment methods.
In the afternoon, when the feature reaches the test environment, Sofi launches the execution agent on the highest-priority cases, while she does exploratory testing on the part that worries her most: what happens when two people try to cancel the same shared booking. That is where she finds a serious bug that no pre-written test case had considered.
At the end of the day, the execution agent has left results with screenshots and three possible bugs. Sofi checks each one: two are real and one is a false positive caused by badly loaded data. She asks the reporting agent to write up the confirmed bugs with the evidence and reviews them before logging them.
What did the agent do? The heavy, repetitive work. What did Sofi do? The decisions, the judgment, the review and the exploratory testing that found the most important bug. That division of labor is, today, the heart of well-done agentic QA.
The trend: why now?
The idea of automating testing with artificial intelligence is not new. For years there have been tools promising "smart" tests. What changed so that agents are now taken seriously? Several factors came together.
Much more capable models
Language models have improved enormously in three skills that are key for testing: multi-step reasoning, using tools correctly and understanding visual interfaces. Not long ago a model would get lost after five or six actions; today the best ones can sustain long tasks, recover from errors and work on the same goal for quite a while.
Standards for connecting tools
Open protocols such as MCP turned connecting an agent to a browser, a repository or a database from an integration project into something you configure in minutes. When connecting tools is easy, many more use cases appear.
More code, faster
Development teams adopted assistants and agents to write code, and that increased the speed and volume of changes. More code in less time means more things to test, and a QA team of the same size cannot keep up using traditional methods. Paradoxically, the AI that speeds up development is one of the reasons quality has become more important than ever: AI-generated code has bugs too, sometimes very subtle ones.
Pressure to release often
With continuous delivery, the testing cycle shrinks. There are no longer weeks of regression before every release. Agents offer a way to maintain coverage without slowing the delivery pace.
Affordable costs
Using advanced models has become much cheaper and more accessible, and today there are agentic tools any QA can use from their own computer without special infrastructure. That democratizes access: it is no longer exclusive to large companies with research teams.
Testing with agents vs. testing agents
When we talk about AI and QA there are two different topics that often get mixed up, and both will be very important in the coming years:
- Testing with agents: using AI agents as a tool to test traditional software. This is what we have been describing throughout the article.
- Testing agents (and AI systems in general): assuring the quality of products that have AI inside, such as chatbots, assistants, recommendation systems or the agents themselves.
This second point opens up an entirely new specialty for QA, with challenges traditional testing did not have:
Non-determinism
A traditional system gives the same output for the same input. A system based on a language model can give a different answer each time. That breaks the classic idea of "expected result = actual result". Instead of comparing exact texts, you have to evaluate whether the answer meets certain criteria: it is correct, it is relevant, it keeps the right tone, it does not make up data, it does not expose sensitive information.
Evaluations (evals)
To measure the quality of an AI system, sets of evaluation cases are used, known as evals, which are run many times to obtain statistical metrics instead of a simple pass or fail. Designing good evals is a lot like designing good test cases: you have to think of representative scenarios, edge cases and clear criteria. It is natural territory for QAs.
One model evaluating another
A widely used technique is LLM-as-a-judge: using a language model to evaluate another model's answers against a rubric. It is practical and scales well, but you have to validate that the judge is reliable by comparing its verdicts with human evaluations. Once again: QA judgment.
Security and adversarial behavior
AI systems have their own vulnerabilities. The best known is prompt injection: malicious instructions hidden in a text, a web page or a document that try to get the model to do something it should not, such as revealing data or performing an unauthorized action. Testing this kind of attack, along with hallucinations, bias and harmful answers, is part of QA work on AI products. If you are interested in security, this is a very interesting intersection between QA and cybersecurity.
Agents that test agents
And yes, agents are already used to test other agents: one agent simulates users with different profiles and intentions, talks to the chatbot under test and records how it responds. It is a way to generate thousands of test conversations that would be impossible to do by hand.
Risks and limitations
It would be dishonest to talk only about the good parts. Agentic QA has real limitations every team should know before adopting it.

Hallucinations and false results
Models can make things up with complete confidence: report a bug that does not exist, claim that a test passed when it did not run it fully, or describe behavior it did not observe. That is why every important result must be backed by verifiable evidence: screenshots, logs, API responses. A report without evidence is not considered valid, whether it comes from a human or an agent.
The oracle problem
To know whether something is right, the agent needs to know how it should be. If the requirements are ambiguous, the agent will fill the gaps with assumptions, and those assumptions may be wrong. An agent can very efficiently validate that the system does what the agent believes it should do, which is not the same as what the business needs.
Tests that always pass
A subtle risk of self-healing: if the agent is free to "fix" failing tests, it may end up adapting the test to the bug instead of reporting it. A test that adjusts itself until it passes tests nothing. Changes to tests must be reviewed and approved by a person.
Reproducibility
Because the agent decides on the fly, two runs of the same goal may follow different paths. That is good for exploring, but bad for regression, where we need to repeat exactly the same thing. That is why many teams use agents to explore and generate tests, and then turn what is valuable into deterministic scripts that run in the pipeline.
Security, permissions and data
An agent with access to tools can do real damage if it makes a mistake or if someone manipulates it: deleting data, changing configurations or sending information where it should not go. Never give an agent access to production or to real customer data without strict controls. You should also check what information is sent to the model provider and whether that complies with company policies.
Costs and time
Every step the agent takes consumes resources. A regression suite run entirely by agents can be slower and more expensive than the same suite in Playwright. Agents do not replace traditional automation everywhere: it pays to use them where they add the most value.
Overconfidence
Perhaps the biggest risk is human: trusting too much. When a tool gets it right many times in a row, we stop checking. And that is exactly when the important error slips through. The QA's critical judgment is, more than ever, the last line of defense.
Good practices for adopting it
If you are thinking about adding agents to your quality process, these practices will save you several headaches:
- Start with low-risk, high-volume tasks. Generating drafts of test cases, test data or reports is ideal to get started. Leave execution on critical environments for later.
- Always keep a human in the loop. Define the points where the agent has to stop and ask for approval: before logging a bug, before modifying a test, before any irreversible action.
- Give it the minimum permissions it needs. If the agent only needs to read, it should not be able to write. If it only needs to test in the QA environment, it should not have credentials for other environments.
- Work in isolated environments. Use test environments, fictitious data and test accounts. Never real customer data.
- Demand evidence. Every result from the agent must come with screenshots, logs or responses that make it possible to verify it.
- Invest in context. Document business rules, conventions and examples of good work. It is the best investment for improving results.
- Version your instructions and configurations. The instructions you give agents are part of your quality process: keep them in the repository, review them and improve them like any other artifact.
- Measure the impact. Compare before and after: test design time, bugs found, false positives, maintenance time. Without data, it is just a feeling.
- Combine approaches. Use agents to explore, design and analyze, and deterministic automation for the regressions that must run the same way every time.
- Train the team. A powerful tool in the hands of someone without testing fundamentals produces a lot of volume and little value.
What changes in the QA role
We come to the question that worries people most: what happens to our jobs? Here is my view after many years in this field.

From executor to strategist
The most mechanical tasks, such as running the same test case for the umpteenth time, transcribing steps or building data spreadsheets, will increasingly be handled by agents. What gains value is what has always been the heart of QA: understanding the business, identifying risks, deciding what to test and what not to, designing a strategy, questioning requirements and having the judgment to know when something is ready to ship.
The fundamentals matter more than ever
To supervise an agent you need to know more than the agent. If you don't know test design techniques, you won't notice that the agent forgot the boundary values. If you don't know how an API works, you won't spot that it validated a response incorrectly. Fundamentals, like those covered by the ISTQBcertification, do not go out of date: they become the foundation for working well with AI.
New skills
- Critical thinking to evaluate what agents produce and detect convincing errors.
- Context and instruction engineering to tell an agent what you need, with which constraints and with which quality criteria.
- Basic technical knowledge of how models, tools and protocols work, to understand what each agent can and cannot do.
- Testing AI systems: evals, non-determinism, security and bias. A specialty with a huge amount of demand ahead.
- Automation: far from disappearing, knowing how to automate lets you combine the best of agents with the best of deterministic scripts.
What about those just starting out?
I often hear the concern that agents will eliminate junior positions. It is true that some entry-level tasks are being automated, but there is also a huge opportunity: someone starting today, with solid fundamentals and a good command of agentic tools, can add value much faster than a few years ago. The key is not to skip the fundamentals to jump straight to the trendy tool. Understanding what a good test case is, how to report a bug and how to think about risk is still the first step. That is why in our free QA course from scratch we start there.
What's coming
Predicting the future in technology is risky, especially in a field that changes every few months. But some trends are already fairly clear.
Agents built into the pipeline
Agents will stop being something the QA uses from their computer and become one more step in the CI/CD cycle: they analyze each change, decide what to test, run exploratory tests on the affected areas and leave comments on the pull request before anyone approves it.
Increasingly specialized teams of agents
Instead of one agent that does everything, we will see teams of agents specialized in performance, security, accessibility, APIs or user experience, coordinated with each other and supervised by QAs who define the overall strategy.
Continuous testing in production
By combining agents with observability, monitoring in production will become smarter: agents that detect anomalous behavior, reproduce the problem in a test environment and propose a regression test so it doesn't happen again. It is Shift-Right taken to another level.
AI quality as a discipline of its own
With more and more products that include AI, testing models and agents will consolidate as a specialty, with its own tools, methodologies and certifications. ISTQB already has AI-related certifications, such as AI Testing, and the field can be expected to keep growing.
Regulation and traceability
Regulatory frameworks such as the European Union's Artificial Intelligence Act are starting to require evaluation, documentation and human oversight for certain AI systems. That will create demand for people able to demonstrate, with evidence, that a system works as it should. Sounds a lot like QA, doesn't it?
What probably won't change
Someone will have to decide what "quality" means for each product, which risks are acceptable and when something is ready for the user. That responsibility cannot be delegated to a tool. Tools change; judgment remains human.
How to start today
If you made it this far and want to move from theory to practice, here is a five-step path:
- Strengthen the fundamentals. If you are starting out, first learn test case design, bug reporting, API testing and automation concepts. Our free course and the ISTQB study material are a good starting point.
- Use AI as an assistant in your daily work. Ask it to review your test cases, suggest scenarios or help you write reports. You will quickly learn what it does well and where it gets things wrong.
- Try an agent on a practice project. Pick one of the websites to practice testing and let an agent analyze a feature, generate test cases and run them. Compare its work with yours.
- Add specialized agents. At QARMY we built k0lmenIA, a free, open source team of AI agents for QA tasks: story analysis, manual and BDD test cases, test data, bug reports, E2E and API execution, and results reports. It is designed for manual QAs and requires no programming knowledge.
- Combine it with automation. When a flow tested by the agent becomes critical, turn it into a stable automated test. For that you can use the k0lmena framework, with Playwright, Cucumber and support for APIs and performance.
And most importantly: share what you learn. This is a new field for everyone, and communities that experiment together learn much faster.
Frequently asked questions
Does agentic QA replace automation with Playwright or Selenium?
No. They are complementary. Agents are very good at exploring, adapting to changes and solving varied tasks, while deterministic scripts are still faster, cheaper and more reliable for regressions that must run the same way every time.
Do I need to know how to program to use QA agents?
For many current tools, no. Some agents are used by chatting in natural language. Still, having technical knowledge will help you understand what the agent does, configure it better and detect when it makes mistakes.
Is it safe to use agents with my company's system?
It can be, if you take precautions: isolated test environments, fictitious data, minimum permissions, human approval for sensitive actions and a review of what information is shared with the model provider. Before using it on a real project, check your company's security policies.
Will AI put QAs out of work?
It will change the work, just as automation did in its day. Mechanical tasks get automated and the demand for judgment, strategy and specialization grows, including testing AI systems. QAs who adapt will have more opportunities, not fewer.
Conclusion
Agentic QA is neither a passing fad nor a magic solution. It is a natural evolution of automation, driven by language models that can reason, use tools and adapt. Used well, it frees quality teams from repetitive work and lets them focus on what really matters: understanding risks, questioning, exploring and deciding.
Used badly, it creates a false sense of security, reports full of noise and tests that pass without testing anything. The difference is made by the judgment of the people who direct it.
So my recommendation is simple: don't be afraid of it, but don't believe everything it says either. Learn the fundamentals, experiment with the tools, measure results and always keep your critical thinking. The future of QA is not humans versus machines; it is humans who know how to direct machines.
If you want to keep learning, join the QARMY WhatsApp channel, where we share news and resources and announce upcoming courses. And if you try agents with your team, tell us how it went.
