Generative AI and Data: rethinking the data lifecycle

Professional working on a laptop with digital cloud, data, and network elements, representing data engineering and GenAI integration across the data lifecycle.

Most conversations about generative AI start with what models can do. Generate content, write code, answer questions or, increasingly, act across systems.

From a data perspective, however, something equally important is happening underneath. Generative AI is not only creating new ways to consume information. It is beginning to change how data solutions are built, tested and operated. At the same time, every step forward in AI reinforces something data teams already know: good outcomes still depend on good foundations.

This creates a two-way relationship. Generative AI and data are becoming increasingly interdependent. AI needs reliable data to create useful results, while AI itself can help teams build and manage data solutions more efficiently.

Better AI still starts with better data

A highly capable model can still produce the wrong answer when the information available to it is incomplete, outdated or inconsistent. This is not a new challenge for data teams, which have spent years dealing with fragmented sources, unclear ownership and different definitions of the same business concept. What GenAI changes is the way those weaknesses become visible.

In a traditional analytics environment, a quality issue may result in a dashboard that does not reconcile or a metric that raises questions. With GenAI, the same problem can result in a fluent, confident answer that appears perfectly reasonable. That makes quality, context and governance even more important, because users may not immediately recognize when the information behind an answer is unreliable.

McKinsey identifies data readiness as one of the main barriers to scaling AI, highlighting challenges around data quality, structure, availability and context. This reinforces a point that is becoming increasingly difficult to ignore: AI ambitions cannot be separated from the maturity of the data environment supporting them.

The same principle is reflected in the way we approach the Data Journey at Xpand IT. Data strategy, engineering, governance, security and scalability create the foundation from which analytics and AI use cases can evolve with confidence. GenAI does not remove the need for those foundations. It increases the importance of getting them right.

GenAI is becoming part of data engineering

The relationship also works in the opposite direction. Generative AI is not only consuming data; it is beginning to support the teams responsible for building and maintaining the systems behind it. In Xpand IT projects, we already consider its use to accelerate data transformations and the development of integration components that bring data in from many different sources, where each API works in its own way and has to be handled separately. It also supports code generation, refactoring and alignment with technical standards.

There is clear productivity value in this, but the bigger benefit is where engineers can redirect their attention. By spending less time on repetitive implementation work, data engineers can focus more on architecture, modelling and the technical decisions that determine whether a solution will remain scalable and maintainable as it grows.

The same applies when teams work with existing systems. GenAI can help them understand unfamiliar code, document transformations and create an initial implementation that can then be reviewed, adjusted and improved by an engineer. It is also very useful when something is not working as expected, helping engineers make sense of a complicated query and find where the wrong result is coming from. This can be particularly useful in complex environments where teams need to move quickly without losing visibility over how the solution is evolving.

The principle remains the same throughout: AI supports the engineering process, but it does not replace engineering judgement. The value comes from combining faster execution with the experience and context needed to make the right technical decisions.

Testing can become broader and faster

Testing is another area where GenAI can change the data lifecycle. Data solutions have to deal with unusual values, unexpected distributions, changes in schemas and failures in upstream systems, making comprehensive testing difficult to achieve manually.

One option is the use of synthetic data. When access to real information is limited or sensitive, AI can help teams create realistic datasets for development and testing without simply reproducing personal data. Within the approach defined for Xpand IT projects, synthetic datasets should remain aligned with the relevant business rules and statistical characteristics of the domain.

GenAI and AI based testing agents can support this process in several ways:

  • Synthetic data to create realistic datasets when access to real information is limited or sensitive;
  • Automated test generation for unit, integration and regression scenarios;
  • Coverage analysis to identify scenarios that have not yet been tested;
  • Test maintenance to keep validation aligned as code and pipelines evolve.

This does not remove the need to decide what really matters to test. It gives teams greater capacity to cover scenarios that delivery pressure might otherwise leave unexplored and helps shorten the feedback loop between development and validation.

Professional working on a laptop, surrounded by digital elements representing data, artificial intelligence, code, security, and data analysis.

Data operations can become more intelligent

Data teams have already automated much of their delivery through CI/CD and DataOps practices. GenAI and AI agents create the possibility of taking that automation further by supporting the configuration and optimization of build, test and deployment pipelines, automating technical validations and helping identify recurring failures.

Consider a failed data pipeline. A traditional monitoring system can tell the team that something broke, while an AI assisted process can potentially help interpret logs, compare the incident with previous failures and suggest where the problem may have originated. The engineer remains responsible for the decision, but reaching that decision can become faster.

This distinction becomes particularly important as organizations explore more autonomous agents. Helping someone diagnose an incident is very different from allowing an agent to modify a production environment without supervision. As automation becomes more capable, the boundaries around what AI can do need to become clearer as well.

More automation requires stronger governance

Greater automation does not reduce the need for governance. It makes governance more operational, because AI systems may now interact directly with data, code, infrastructure and business processes.

The approach defined for GenAI in Xpand IT projects establishes several principles that become increasingly relevant as these capabilities evolve:

  • Control over the data used in prompts and models;
  • Security, confidentiality and compliance throughout the lifecycle;
  • Traceability over AI generated outputs and actions;
  • Mandatory human review for critical code, business logic and technical decisions.

There is a practical reason for these safeguards. Generated code can be technically correct and still apply the wrong business rule, while a testing agent can create extensive coverage and still overlook the scenario carrying the greatest business risk. An automated recommendation can also make sense based on the information available to it while missing context that only the team understands.

Human review is therefore not simply a temporary measure while models improve. It is part of responsible engineering. The objective is to take advantage of the speed AI can provide without losing the accountability and context required to make that speed sustainable.

The data lifecycle is evolving

It would be easy to describe all of this purely as a productivity story. GenAI can help engineers write code faster, create more tests and investigate incidents more efficiently, and those gains are valuable. The bigger opportunity, however, lies in what teams can do with the capacity they recover.

Less repetitive implementation work can mean more attention to architecture and data quality. Better testing can make experimentation safer, while faster diagnosis can give engineers more time to improve platforms instead of repeatedly resolving the same operational problems. Together, these changes can shorten the distance between a business requirement and a reliable data product. 

The models and tools behind these capabilities will continue to evolve quickly, which is precisely why a data strategy should not be built around a specific AI product. The more durable capabilities remain familiar: reliable data, clear ownership, scalable architecture, strong engineering, testing, security and governance.

GenAI is changing how teams work across all of them. It is not replacing the data lifecycle; it is becoming part of it.

Search

Most Popular