Skip to main content

Why Research Workflows Still Matter in the Age of AI Agents

Gatsbi

More autonomy is not always the same as more value. For research, the quality of the process matters as much as the capability of the model.

Why use a workflow-based AI research assistant when a general-purpose agent could take a research goal and decide how to pursue it?

It is a fair question. It deserves a better answer than claiming that agents are unreliable, or that every research problem should follow a fixed sequence.

At Gatsbi, our answer starts with a distinction: being able to perform individual research tasks is not the same as providing a dependable research experience.

A researcher needs more than an AI that can produce an impressive response. They need a useful connection between their question, the evidence, the analysis, and the document they eventually review and share. They need to know where their judgment matters, what the system has produced, and what still requires verification.

That is where a workflow-based AI research assistant creates its value—not by limiting intelligence, but by organising it around the work researchers actually need to complete.

Workflows and Agents Are Not Opposites

The distinction is less about whether a system is “intelligent” and more about who controls the process.

In a workflow-led system, the application defines more of the sequence and the boundaries between tasks. In a more open-ended agent system, the model has greater freedom to decide what to do next and which tools to use. These approaches can coexist: a structured workflow can contain autonomous research, iterative reasoning, and tool-using agents within its stages.

Gatsbi already reflects this combination. Its research workflows include an integrated Deep Research Agent for collecting and synthesising evidence, alongside dedicated experiences for research ideation, manuscript drafting, systematic reviews, and meta-analyses. Calling Gatsbi workflow-based does not mean it is simply a static sequence of prompts. It describes how these capabilities are organised into a research product.

The useful question is therefore not “workflow or agent?” It is:

Which decisions benefit from flexibility, and which should remain explicit, consistent, and under the researcher’s control?

Where Gatsbi’s Workflow-Based Approach Creates Value

Research expertise is built into the process

A research assistant should contribute more than general language ability. It should help organise the work according to the requirements of the task.

Consider a systematic review. Finding relevant papers and writing a fluent summary are only parts of the job. A rigorous review also depends on a clearly defined question, explicit eligibility criteria, appropriate methods, and careful handling of the evidence. The Cochrane Handbook emphasises prespecified methods, protocols, and quality assurance precisely because the process affects the reliability of the conclusions.

This is a strong reason to use purpose-built workflows. Instead of asking the model to invent a research procedure from a broad instruction, the product can make the expected stages part of the experience.

Gatsbi Reviewer, for example, connects study identification and screening with data extraction, synthesis, review of the results, and manuscript generation. These are presented as connected research activities rather than unrelated requests in a conversation.

The advantage is not that a workflow automatically makes the research correct. It is that important methodological steps do not have to depend entirely on the user remembering to request them.

For researchers, that changes the starting point. They can spend more attention evaluating the question and the evidence, rather than repeatedly explaining how a research process should be organised.

Researchers spend less effort managing the AI

There is a difference between delegating a task and becoming the project manager for a collection of AI interactions.

Imagine working through a literature-based project using a general-purpose agent. Even with a capable system, someone must decide what the deliverable should contain, establish the research scope, explain the relevant methods, review intermediate work, and determine when the result is ready to move forward.

A well-configured agent can handle much of this. But configuring, testing, and maintaining that setup is itself work.

Gatsbi’s proposition is to make the common path available as a product. For projects that fit its supported workflows, users do not have to design every handoff between research, analysis, writing, and export.

This matters particularly when someone wants to complete a familiar research task rather than develop a custom automation system. A researcher preparing another evidence synthesis may value a clear, reusable process more than the freedom to redesign the entire process on every occasion.

A sufficiently engineered general-purpose agent could reproduce many specialised functions. That does not make the specialised product unnecessary. “Possible to assemble” and “ready to use” are different kinds of value.

Human control is placed where it matters

There is an important difference between being allowed to interrupt an AI and being given a meaningful opportunity to review its work.

For research, useful control should happen before an upstream mistake becomes embedded in a polished manuscript.

Gatsbi Reviewer makes this concrete. Users can curate the included studies and review or edit extracted data, statistical outputs, visual summaries, and synthesised findings before generating the manuscript.

The significance is not simply that there is an edit button. It is where that editing happens.

Imagine that a study has been assigned to the wrong category, or that an extracted value needs correction. Reviewing the relevant material before drafting gives the researcher an opportunity to address the issue closer to its source, rather than discovering it after it has influenced several sections of the paper.

This is a useful principle for research automation: the system should not ask the user to choose between doing everything manually and accepting everything automatically.

A workflow can provide a middle ground—automating the repetitive work while making consequential decisions visible and reviewable.

The destination is a research artifact, not just a conversation

A research task does not end when the AI has something convincing to say. It ends with material that can be inspected, revised, discussed, and incorporated into further work.

Gatsbi Writer is designed around that destination. It supports manuscript drafts with citations, references, equations, figures, and tables, with export to Word, LaTeX, and Markdown. It also uses different writing workflows for different research types, rather than treating every project as the same generic essay.

The practical value lies in the connections between these elements.

A useful manuscript requires more than individually good paragraphs. The methods must correspond to the work described. Tables and figures must support the discussion. References must remain connected to the claims they are intended to support. The document must also be editable in the environment where the researcher continues working.

A general-purpose agent can produce documents too. Gatsbi’s advantage is that research-oriented deliverables are part of the product’s default objective, rather than an output specification that the user must reconstruct each time.

This does not make the draft submission-ready without review. It makes the draft a more useful starting point for that review.

A defined process creates a clearer path to consistency and improvement

For repeatable tasks, a defined workflow gives developers something concrete to test.

Instead of evaluating only whether the final manuscript “looks good,” a workflow-based system can be assessed at specific stages: whether an extraction contains the required fields, whether a calculation uses the intended inputs, or whether the generated document contains the expected components.

This is an architectural opportunity, not a guarantee that every possible check has already been implemented. Its value is that failures can be investigated at the level where they occur.

Workflow decomposition also makes it possible to give individual model calls narrower responsibilities and place programmatic checks between them. Engineering guidance on agentic systems identifies this as a useful approach for tasks that can be divided into well-defined subtasks.

For users, the intended benefit is consistency of process—not identical wording on every run, and certainly not guaranteed scientific correctness.

The same distinction applies to cost. A predefined process creates opportunities to limit unnecessary exploration, reuse components, and avoid repeatedly planning familiar steps. But a badly designed workflow can still be slow or expensive.

The objective should be a useful result with proportionate effort, not the largest possible number of autonomous actions.

The Limitations Deserve an Honest Assessment

The strengths of workflows come with trade-offs.

A predefined process cannot anticipate every research problem. An exploratory project may change direction after a surprising finding. An interdisciplinary study may require an unusual combination of methods. A researcher may need to connect a private database, execute a custom analysis, or work with a tool that the product does not support. In these situations, a more open-ended agent—or a custom research environment—may be the better fit.

There is also a risk of confusing procedural completeness with substantive quality. A system can complete every stage and still work from incomplete evidence, misunderstand a source, or produce an analysis whose assumptions are inappropriate. A well-structured manuscript does not establish that its conclusions are justified.

That is why research expertise remains essential. In systematic reviewing, for example, established guidance treats methodological expertise, domain expertise, and quality assurance as central requirements, not optional additions that a completed checklist can replace.

Nor should an AI-generated research draft be treated as evidence that an experiment occurred or that a proposed idea is genuinely novel. Gatsbi’s own guidance calls for source verification, revision, and careful attention to originality and academic integrity.

Finally, workflow products must keep earning their place. General-purpose agents can also use structured procedures, specialised tools, and review checkpoints. The distinction is not a permanent technical boundary.

Gatsbi’s value must therefore come from how well its research experience is designed, integrated, and maintained—not from claiming that an agent could never perform the same tasks.

What We Are Building Next: Flexibility Without Losing Control

These trade-offs are helping shape our next product.

Our team is developing an agent-based product that will combine the advantages of structured workflows with the flexibility of Skills: reusable packages of task-specific instructions, procedures, and resources that an agent can draw on when needed. This modular approach makes it possible to provide specialised capabilities without placing every instruction into every task’s context.

Our intended division of responsibilities is straightforward. Workflows will provide the structure for activities that benefit from a dependable process. Skills will supply specialised knowledge and methods. The agent will have room to select and combine capabilities when a task requires a less predictable path.

Crucially, instructions alone are not the same as enforced controls. The goal is to preserve meaningful boundaries and review points while allowing flexibility within them.

We are also developing our own agent harness—the execution and control layer around the model—to help smaller models work more reliably. Harness design addresses issues such as maintaining task state, managing context, checking progress, and recovering from failures; published work on long-running agents illustrates why these surrounding mechanisms matter alongside model capability.

For us, this work has an important economic purpose: reducing token costs without making researchers absorb the cost of unreliable execution.

A lower-priced model is not genuinely cheaper if repeated failures, unnecessary retries, and manual corrections erase the savings. Our aim is to make smaller models dependable for suitable tasks through better system design, rather than assuming that every step requires the largest available model.

We will judge that approach by reliability and total task cost—not by model size alone. These are development goals, not a claim of a benchmarked cost reduction already achieved.

The next product is therefore an extension of the reasoning behind Gatsbi, not a rejection of it.

Research needs flexibility when the path is uncertain, structure when the method matters, and human judgment where the conclusions carry weight. The opportunity is to bring those strengths together—making AI research assistance more capable, more controllable, and more affordable.