THE POLICY EDGE
Reports/Data Releases

13 August 2026

Parliamentary Panel Proposes AI Roadmap for Gyan Bharatam Mission

The Mission has built a database of more than 1.19 crore manuscripts and digitised about eight lakh, but the committee finds that converting these images into accurate, searchable and research-ready texts will require better capture standards, large-scale human verification and safeguards for custodian consent

Listen to the article
Reports/Data Releases image

Key Details

The Gyan Bharatam Mission, launched to survey, document, conserve and provide access to India’s manuscript heritage, has an approved outlay of ₹491.66 crore for 2025–2031. The 393rd Report of the Parliamentary Standing Committee on Transport, Tourism and Culture examines how AI can help the Mission move from large-scale cataloguing and scanning to accurate transcription, research use and long-term preservation.

Mission Progress

What the Committee Says Comes Next

More than 1.19 crore manuscripts reported

Prioritise manuscripts by condition and scholarly value

About 8 lakh manuscripts digitised

Capture faded writing, palm-leaf inscriptions and material information that ordinary scans may miss

Handwriting recognition reached 90–91% accuracy

Build a citizen-and-university workforce to verify machine-generated text

Only about 250 manuscripts completed the full processing pipeline

Measure progress through verified, usable texts—not scans alone

More than 3.90 lakh manuscripts publicly viewable

Introduce custodian-controlled access and labels identifying unverified AI output

72 GPUs provided through MeitY

Develop open Indic-script training data and allow models to learn from restricted collections without removing their images


Digitisation Has Created a Verification Bottleneck

The Gyan Bharatam Mission has progressed rapidly in cataloguing and scanning manuscripts. The next challenge is turning those images into dependable texts.

AI tools can recognise handwritten scripts, but reported accuracy of 90–91% still leaves errors that can alter names, dates, formulations or meaning. Human verification therefore remains necessary before machine-read texts can support translation, research or public use.

The Committee proposes:

  • an Akshar Mitra platform where two volunteers independently verify each page, with disagreements referred to another contributor or scholar; and

  • university partnerships allowing students to prepare critical editions of unedited manuscripts as dissertation and research work.

The aim is to expand verification beyond a small pool of manuscript specialists.


Better Imaging Could Recover Otherwise Lost Information

Ordinary photography may miss faded writing, overwritten text or inscriptions incised into palm leaves. Since many manuscripts may be handled only once, the Committee proposes a richer national capture standard:

  • multispectral imaging for faded, damaged or overwritten pages;

  • surface-relief imaging for palm-leaf manuscripts;

  • non-destructive analysis of ink, paper, bark and leaf; and

  • scientifically anchored dating with uncertainty clearly disclosed.

A separate research challenge would develop methods for reading manuscripts that are fused, burnt or too fragile to open.

The principle is “capture once, capture completely”: information missed during mass digitisation may be costly or impossible to recover later.


AI Outputs Would Carry Verification Labels

The Committee recommends a persistent credential for every AI-generated transcription, translation and audio rendering, identifying:

  • the source;

  • creation date;

  • machine generation; and

  • scholarly verification status.

This would distinguish unverified AI output from authoritative text as material is downloaded and shared.

Private custodians would decide, manuscript by manuscript, whether an item may be copied, only viewed, restricted to a community or kept private. For collections whose owners do not permit images to leave their premises, the committee recommends federated learning, allowing AI models to improve without transferring the underlying manuscript images. This is important because the Mission depends on an estimated 16,000–17,000 private manuscript holders.


The Corpus Could Become a Research Dataset

Once manuscripts are accurately transcribed, AI could connect people, places, dates and works across languages and collections.

The Committee proposes:

  • an open index of persons, places and works;

  • a time-layered atlas of knowledge production and transmission;

  • historical datasets on eclipses, weather, floods and famines;

  • digital reconstruction of manuscripts dispersed across collections; and

  • links between traditional medical texts and the official pharmacopoeia, while retaining approved standards for medical practice.

More than two lakh Indian manuscripts across 54 institutions in 23 countries have also been identified. The Committee seeks digital copies, virtual reconstruction of dispersed manuscripts and negotiated return of originals where possible.


Cultural Mapping Would Support Planning and Livelihoods

The recommendations extend to oral traditions, endangered languages and vanishing art forms. The Committee proposes:

  • prioritising languages and oral traditions with the fewest surviving bearers;

  • free fonts, keyboards and computer voices for inadequately supported Indian scripts;

  • 3D and motion recording of traditional practitioners’ techniques; and

  • connecting mapped artisans with buyers through ONDC.

The cultural map, covering 97.7% of identified villages, would also feed into PM Gati Shakti and district planning, allowing cultural assets to be considered before infrastructure alignments are finalised.


Policy Relevance

  • Digitisation targets should distinguish between scans and usable knowledge. Manuscripts photographed, machine-read, human-verified, translated and publicly released are different stages and should be reported separately.

  • Capture standards are time-sensitive. Multispectral and surface-relief imaging must be incorporated before large-scale contracts finish, not added after fragile manuscripts have been returned.

  • Public participation needs quality assurance. Independent checks, contributor training and expert escalation will determine whether Akshar Mitra improves accuracy at scale.

  • Custodian consent will shape the size of the national collection. Enforceable access choices and secure learning methods can encourage private holders to participate without surrendering control.

  • AI provenance should become a standard for public digital archives. Every machine-generated cultural record should disclose its source and verification status.

  • The Mission’s value will ultimately be judged by use. Research datasets, critical editions, scientific discoveries and livelihood opportunities provide a stronger measure of return than the number of pages scanned.


Follow the Full Report Here: From Manuscript to Mission: Preserving, Mapping and Sustaining India’s Cultural Heritage in the Age of Artificial Intelligence—393rd Report of the Department-related Parliamentary Standing Committee on Transport, Tourism and Culture

Rethinking Public Policy Through Insight | Inquiry | Impact

Opinion • Grassroots Voices • Policymakers Perspectives • Expert Analysis • Policy Briefs