AI Modernization
Independent analysis of AI-assisted migration tools, COBOL-to-Java translation benchmarks, AI readiness blockers, and MLOps cost data.
On this page
- Overview
- Direction A: AI as a Modernization Tool
- AI Code Translation: What the Benchmarks Actually Show
- AI Test Generation for Legacy Code: Documented Results
- AI Legacy Code Documentation Tools
- Direction B: Modernization for AI Readiness
- The Blockers: Why 60% of Companies See No AI Value
- What AI-Ready Data Infrastructure Looks Like
- MLOps Maturity: The Gap Between Pilot and Production
- AI Readiness Modernization Path: Three Phases
- Cost Benchmarks
- Testing Cost Transformation: Manual → AI-Native (Virtuoso 2025)
Two directions, one dependency. AI tools are becoming useful for code translation, test generation, and documenting legacy systems, and vendors and researchers now publish measurable results for each. But AI projects stall when data is fragmented and legacy integration is hard: in Cisco's 2025 AI Readiness Index, only 19% of organizations had fully centralized data infrastructure (as reported by RCR Wireless). This hub covers both.
Important
Two Directions, One Research Hub
This hub covers a two-way relationship between AI and modernization that is often conflated into one narrative. We separate it into two questions:
| Direction | Focus | What It Covers |
|---|---|---|
| Direction A | AI as a Modernization Tool | Using LLMs and AI tools to accelerate legacy code migration, generate test suites, and document undocumented systems. The tooling is real and maturing fast. |
| Direction B | Modernization for AI Readiness | What organizations must modernize — data infrastructure, APIs, MLOps pipelines — before AI can deliver value in production. Most companies are blocked here. |
AI-assisted code translation has become practically useful. IBM Research reports a median structural quality score above 80% for its watsonx Code Assistant for Z pipeline, measured on a vendor-run evaluation of 20 sample programs and 10 customer applications (IBM Research), and one testing vendor's own ROI model claims AI-native testing costs a fraction of manual testing for a hypothetical company (see Cost Benchmarks). At the same time, BCG's 2025 survey found that only 5% of firms are "future-built" for AI and achieving value at scale, while 60% report minimal revenue and cost gains despite substantial investment (BCG). BCG does not attribute that gap to data infrastructure alone, but data fragmentation and legacy integration are among the blockers covered below.
Read full background
The AI tooling market for software modernization has bifurcated. General-purpose tools (GitHub Copilot, Google Gemini Code Assist) are built for everyday developer assistance; no public benchmark of them on enterprise COBOL translation was found for this review. COBOL's data division structure, packed-decimal arithmetic, and CICS transaction context are where generic translation is most likely to need close review. Specialized tools (IBM WCA4Z, AWS Transform for mainframe) are purpose-built for legacy modernization, and IBM publishes evaluation results for its own pipeline; validate any tool on a sample of your own code.
On the readiness side, Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, or inadequate risk controls (Gartner). Deloitte separately names legacy system integration as one of three infrastructure obstacles to agentic AI (Deloitte). Pilots tend to succeed in sandboxed environments with clean, curated data. Production deployment gets harder when the agent must call a batch system via SFTP, read from several inconsistent CRMs, or wait for a nightly ETL job to refresh the data it needs in real time. Modernization is often a prerequisite for AI value, not an alternative to it.
Direction A: AI as a Modernization Tool
AI tools are demonstrably useful for three specific modernization tasks: code translation, test generation, and documentation of undocumented systems.
AI Code Translation: What the Benchmarks Actually Show
IBM watsonx Code Assistant for Z
IBM Research reports a median structural quality score above 80% and a functional score above 75% for its translation pipeline, evaluated on 20 GenAPP sample programs and 10 proprietary customer applications (IBM Research). These are vendor-reported results on a small evaluation set. A separate IBM Research study found that adding natural-language summaries to COBOL source before translation improved 36% of eligible CodeNet samples (727-sample benchmark) and 50% of low-scoring samples on an enterprise benchmark (571 samples) (IBM Research). Validate any figure on a sample of your own code.
A specialized option for COBOL-to-Java; validate on your own code
GitHub Copilot & General-Purpose LLMs
Built for everyday developer assistance and greenfield code. IBM's evaluation compared its own pipeline with three standalone LLMs, but no independent public benchmark of GitHub Copilot on enterprise COBOL translation was found for this review. COBOL data structures (COMP-3 packed decimal, REDEFINES, 88-level condition names), database handling, and CICS context are where generic translation needs the closest review. Pilot it on a COBOL sample before relying on it for legacy migration.
Use for modern code; pilot on your own COBOL sample before relying on it for legacy migration
| TOOL | BEST USE CASE | COBOL EVIDENCE | PRICING |
|---|---|---|---|
| IBM watsonx Code Asst. for Z | COBOL → Java, mainframe migration | IBM-reported: 80%+ structural, 75%+ functional (20 sample programs and 10 customer applications) | Enterprise contract |
| Amazon Q Developer / AWS Transform | Java upgrades, .NET migration, mainframe refactor | No public COBOL benchmark found | Q Developer IDE subscriptions reach end of support 30 April 2027; code transformation moves to AWS Transform (AWS) |
| GitHub Copilot | Daily coding assistance, greenfield | No public COBOL benchmark found | Per-seat plans with usage-based billing; see GitHub |
| Google Gemini Code Assist | GCP workloads, modern languages | No published legacy benchmarks | Per-seat subscription; see Google Cloud |
| Moderne | Large-scale Java/Spring refactoring | Java-focused; not a COBOL tool | Enterprise contract |
AI Test Generation for Legacy Code: Documented Results
UnitTenX (2025, authors' results — arXiv)
100%
line coverage reported on one real-world C codebase starting from 0%
coverage measured for 186 of 199 executed functions (93.5%)
Salesforce / Cursor AI
85%
reduction in time to reach legacy code coverage targets
26 → 4 engineer days per module; some repos <10% initial coverage (Salesforce Engineering)
Virtuoso (2025, vendor ROI model — source)
$840K
annual testing cost with AI-native platform, in the vendor's model
vs $6.03M manual, $2.29M traditional automation (hypothetical mid-market SaaS company)
Where AI Test Generation Fails
AI-generated tests cover code paths reliably but cannot validate that the business rules encoded in those paths are correct. For undocumented legacy systems, the business logic is the unknown — the AI has no way to verify that a 1998 calculation produces the right answer because there is no specification to check against. The Salesforce approach — mandatory human review of all generated tests, AI-generated JavaDoc comments for human validation, SonarQube as final gate — is the right model. Treat AI test generation as a way to get coverage, not to verify correctness.
AI Legacy Code Documentation Tools
LegacyMap (Sector7)
Specialized for COBOL, FORTRAN, BASIC, C++, and Pascal on OpenVMS and Mainframe (z/OS). Generates structured callgraphs, SQL access diagrams, and procedure maps without requiring code changes. Supports platform-specific dialects including VAX, Alpha, and z/OS variants. Best tool for generating structural maps of undocumented mainframe estates — accelerates onboarding and identifies dead code without modifying production systems.
Codegram
Targets VB, Delphi, and COBOL to modern languages (Java, C#, Python) with integrated documentation generation, debugging tools, and code optimization. Provides structured documentation templates and maintains conversion history for auditing. Combined conversion + documentation toolchain for mid-market legacy application migrations. Appropriate for organizations that want a unified tool rather than separate documentation and translation workflows.
Quality ceiling on documentation tools: All documentation tools excel at structural analysis — identifying callers/callees, visualizing system flows, and accelerating engineer onboarding. None can reconstruct business intent from undocumented logic. They map what the code does, not what it was supposed to do or why specific decisions were made. Human domain experts remain required for semantic validation.
Direction B: Modernization for AI Readiness
Before AI delivers value in production, the underlying infrastructure must support it. Most enterprises are blocked at data and integration layers.
The Blockers: Why 60% of Companies See No AI Value
Data Fragmentation
Cisco's 2025 AI Readiness Index, a global survey of 8,000 senior IT and business leaders, found that only 19% of organizations have fully centralized data infrastructure (as reported by RCR Wireless). A typical failure: siloed CRM, POS, ecommerce, service, and transaction data that cannot be joined, which makes it hard to calculate customer lifetime value or give AI agents full context.
Legacy System Integration
Deloitte's Tech Trends 2026 lists legacy system integration among three infrastructure obstacles to agentic AI, noting that agents mostly rely on APIs and conventional data pipelines to reach enterprise systems (Deloitte). Gartner predicts over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, or inadequate risk controls (Gartner). A common failure pattern: AI pilots succeed in isolated environments with clean test data, then production deployment stalls when connecting to legacy systems without APIs, real-time data, or the query patterns that AI agents generate. Monolithic architectures, tightly coupled systems with brittle integrations, and slow release cycles are typical contributors.
Important
The AI Value Gap (BCG and McKinsey, 2025 surveys)
| Metric | Value | Context |
|---|---|---|
| AI future-built companies | ~5× revenue gains | vs AI laggards, ~3× cost reductions (BCG) |
| Share of firms that are future-built | 5% globally | 35% scaling; 60% reaping minimal value |
| Respondents attributing any enterprise-level EBIT impact to AI | 39% | McKinsey State of AI 2025; most of those say under 5% of EBIT (McKinsey) |
What AI-Ready Data Infrastructure Looks Like
AI-Ready Enterprise
- ✓API integration layers over core systems
- ✓Data lake/lakehouse architecture
- ✓Enterprise data warehouse for structured analytics
- ✓Data governance with contracts, freshness SLAs, and error budgets
- ✓Event-driven pipelines that activate instantly on business events
- ✓Federated data mesh: teams publish data products under shared contracts
- ✓Policy-as-code: consent, residency, and retention in version control
Legacy Enterprise (AI Blocked)
- ✗Siloed CRM, POS, ecommerce, service data — no joins possible
- ✗No APIs on core systems — batch SFTP or database polling only
- ✗Nightly ETL — agents can't get real-time data
- ✗"One-off" integrations with no reusable contracts or schemas
- ✗No data lineage or quality controls — models train on stale data
- ✗Manual approvals for deploying or retraining models
- ✗Inflexible permission structures blocking experimentation
MLOps Maturity: The Gap Between Pilot and Production
The gaps below are common reasons AI does not reach production at legacy organizations. Readiness checklists typically ask for a data-quality baseline (this page uses 80%+ as an editorial target), executive sponsorship, cloud ML infrastructure, and defined ROI metrics before model deployment.
No MLOps CI/CD Pipeline
Models are trained by data scientists but deployed manually — inconsistently and infrequently. No automation for model validation, staging, rollback, or A/B testing. Each deployment is a one-off project.
No Production Monitoring
Models degrade silently when production data drifts from training data. No alerting when model performance drops. No feedback loops into retraining cycles. Models deployed in 2024 still running without updates in 2026.
Undefined Roles
Organizations have data scientists but no ML engineers. No one owns the boundary between model development and production deployment. MLOps responsibilities fall into a gap between data science and IT operations.
No Data Governance
Models trained on data with no quality controls, freshness guarantees, or bias audits. Regulatory exposure increases as AI decisions become consequential. EU AI Act compliance requires documentation that doesn't exist.
AI Readiness Modernization Path: Three Phases
Phase 1 — 3–6 Months
Data Foundation
- Data quality audit (target 80%+ score)
- API layer over legacy core systems
- Master data governance baseline
- Cloud infrastructure rebalancing
- Data sovereignty compliance mapping
Phase 2 — 6–12 Months
Integration & Pipelines
- Data lakehouse implementation
- Real-time event-driven pipelines
- MLOps CI/CD pipeline build
- Model monitoring & observability
- Feature store for shared ML features
Phase 3 — 12–24 Months
Production AI
- Federated data mesh with guardrails
- Policy-as-code for AI governance
- Fine-tuning on proprietary data
- Agentic system deployment
- EU AI Act / regulatory compliance
Phase durations are editorial planning ranges, not measured averages. Cisco's 2025 AI Readiness Index found that 13% of surveyed organizations reached its top readiness tier, which it calls Pacesetters (as reported by RCR Wireless). Timeline compresses when data quality is already high; extends when core systems are on-premise with no APIs.
Cost Benchmarks
A vendor-published 2025 cost comparison for testing. It is a vendor model, not an independent measurement.
Testing Cost Transformation: Manual → AI-Native (Virtuoso 2025)
| APPROACH | ANNUAL COST | VS MANUAL | 3-YEAR SAVINGS |
|---|---|---|---|
| Manual Testing | $6,030,000 | Baseline | — |
| Traditional Test Automation (Selenium/Cypress) | $2,290,000 | −62% | $11.2M |
| AI-Native Testing | $840,000 | −86% | $15.6M |
- Source: Virtuoso QA ROI calculator, 2025. Vendor model for a hypothetical mid-market SaaS company, not measured customer results. 3-year savings vs manual baseline. Does not include the competitive advantage value ($47M+) that Virtuoso estimates for faster release velocity.
AI Modernization Services & Vendor Guide
Compare 10 AI modernization tools and strategy partners, see adoption data, and explore service offerings.
Frequently Asked Questions
How accurate is AI code translation from COBOL to Java in 2026?
IBM Research reports that its watsonx Code Assistant for Z (WCA4Z) translation pipeline reaches a median structural quality score above 80% and a functional score above 75%, evaluated on 20 sample programs and 10 customer applications against three standalone LLMs ([IBM Research](https://research.ibm.com/publications/enterprise-scale-cobol-to-java-translation-llms-augmented-with-program-analysis)). These are vendor-reported results on a small evaluation set. A separate IBM Research study found that adding natural-language summaries to COBOL source before translation improved 36% of eligible samples on a 727-sample open benchmark and 50% of low-scoring samples on a 571-sample enterprise benchmark ([IBM Research](https://research.ibm.com/publications/refining-llm-based-cobol-to-java-translation-via-natural-language-summary-augmentation)). No independent public benchmark of GitHub Copilot on enterprise COBOL was found for this review, so claims either way are unverified. Manual refinement is reduced but not eliminated; plan for substantial human review of any enterprise COBOL migration and validate every tool on a sample of your own code.
What are the real hallucination rates for AI code translation tools?
No authoritative hallucination or error rate for AI code translation of legacy languages is published here. Most public hallucination figures come from general-language benchmarks, such as summarization leaderboards, and do not transfer to COBOL-to-Java translation. Treat the rate as something to measure: run the tool on a representative sample of your own code and compare the output with the original behavior. For production code, automated evaluation is required, and so is manual review by engineers who understand both the source language semantics and the target language idioms. Do not deploy AI-translated code without a regression test suite.
Can AI generate test suites for legacy code that has no existing tests?
Yes, with documented results. The authors of UnitTenX (an open-source AI multi-agent system, 2025) report taking one real-world C codebase from 0% to 100% line coverage, with coverage successfully measured for 186 of the 199 functions that executed (93.5%) ([arXiv](https://arxiv.org/abs/2510.05441)). That is one codebase, reported by the tool's authors. Salesforce's Einstein Activity Capture team reported that using Cursor cut unit-test effort from 26 engineer days to 4 per repository module, an 85% reduction, on legacy repositories where some had less than 10% initial coverage ([Salesforce Engineering](https://engineering.salesforce.com/how-cursor-ai-cut-legacy-code-coverage-time-by-85/)). The Salesforce approach combined AI generation with mandatory human oversight: engineers reviewed all generated code, checking test intentions through AI-generated JavaDoc comments, with SonarQube as the final validation gate. Where AI-generated tests fail: business logic validation in undocumented legacy systems — the AI can cover code paths but cannot verify that the business rules encoded in those paths are correct.
What are the biggest blockers preventing enterprises from adopting AI in 2026?
Surveys point to data readiness and legacy integration as leading blockers, but rankings differ by survey and no single study settles which is first. Cisco's 2025 AI Readiness Index found that only 19% of organizations have fully centralized data infrastructure ([as reported by RCR Wireless](https://www.rcrwireless.com/20251209/ai/cisco-ai-readiness-index-2025)). Deloitte's Tech Trends 2026 names legacy system integration as one of three infrastructure obstacles to agentic AI ([Deloitte](https://www.deloitte.com/us/en/insights/topics/technology-management/tech-trends/2026/agentic-ai-strategy.html)). Gartner forecasts that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, or inadequate risk controls ([Gartner](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027)). A common failure pattern: AI pilots succeed in isolated environments with clean test data, then stall when connecting to legacy systems without APIs or real-time data capabilities.
What does 'AI-ready data infrastructure' actually look like?
AI-ready estates typically combine three core layers: API integration, a data lake or lakehouse, and an enterprise data warehouse (no adoption percentages are published here). Beyond technology, AI-ready infrastructure requires: data governance with operational data contracts, freshness SLAs, and error budgets; real-time event-driven pipelines that activate instantly on business events; federated data mesh with centralized guardrails where regional teams publish data products under shared contracts; and policy-as-code embedding consent, residency, and retention rules. Legacy infrastructure fails on all of these: siloed CRM, POS, ecommerce, and service data creating blind spots, no lineage or quality controls, and 'one-off' integrations lacking reusable contracts.
What is the cost of NOT modernizing for AI — what are companies actually losing?
BCG's 2025 survey of more than 1,250 executives reports that 'future-built' AI companies achieve about 5x the revenue increases and 3x the cost reductions of laggards. Only 5% of firms qualify as future-built, 35% are scaling, and 60% report minimal revenue and cost gains despite substantial investment ([BCG](https://www.bcg.com/publications/2025/are-you-generating-value-from-ai-the-widening-gap)). McKinsey's State of AI 2025 survey found that only 39% of respondents attribute any enterprise-level EBIT impact to AI, and most of those say it is under 5% of EBIT ([McKinsey](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai)). These are survey results, not a measured cost of delay for your business; build that estimate from your own use cases.
How long does an AI readiness transformation actually take?
The typical AI readiness modernization path has three phases: Phase 1 (3–6 months): data audit, master data governance, API layer over legacy systems, cloud rebalancing to get workloads on infrastructure that supports AI inference; Phase 2 (6–12 months): data lakehouse implementation, real-time event pipelines, MLOps CI/CD pipeline build, model monitoring infrastructure; Phase 3 (12–24 months): federated data mesh, policy-as-code, retraining infrastructure on proprietary data, agentic system deployment. The timeline compresses when data quality is already high (rare) and extends when ERP/CRM systems are on-premise with no APIs (common). For context, Cisco's 2025 AI Readiness Index found that 13% of surveyed organizations reached its top readiness tier, which it calls Pacesetters ([as reported by RCR Wireless](https://www.rcrwireless.com/20251209/ai/cisco-ai-readiness-index-2025)). The phase durations above are editorial planning ranges.
What are the MLOps maturity gaps that prevent AI from reaching production?
Common organizational gaps include: lack of defined roles and skill requirements for MLOps (teams have data scientists but no ML engineers), insufficient CI/CD pipelines for ML model deployment (models are trained but deployment is manual), inadequate monitoring and observability for production models (no alerting when model performance degrades), absence of data governance frameworks (models trained on stale or biased data), and no clear resource allocation for MLOps at various maturity levels. Readiness checklists commonly require a data-quality baseline (this page uses 80%+ as an editorial target), executive sponsorship, and defined ROI metrics before model deployment.
Chief Analyst, Software Modernization Intelligence · 10+ years B2B market research
Last reviewed: