Success Story
About The Client

IndiaAI, under the Ministry of IT, drives the country’s AI initiatives through a seven-pillar strategy. In collaboration with the National e-Governance Division (NeGD), it launched AIKosh, India’s first national AI datasets platform, centralizing datasets, models, and resources to provide a trusted hub for AI development and training.
Business Situation
IndiaAI’s motive was to consolidate fragmented datasets across ministries, universities, hospitals, and private organizations and develop a data repository to train AI models. Data engineers spent hours sourcing, cleaning, and validating information from multiple sources. Private datasets were costly, inconsistent, and lacked standardized licensing frameworks. Limited access to GPUs further slowed AI model training, making development expensive, time-consuming, and unreliable.
IndiaAI wanted a single national repository consolidating datasets from government and private contributors. The platform had to provide affordable, authentic, and privacy-compliant data. It also needed built-in tools for experimentation and model training, with the ability to scale and onboard more ministries, organizations, and contributors. Unthinkable developed this solution.
Key requirements were:
Conduct an initial discovery phase to identify fragmented datasets, stakeholders, and integration requirements.
Aggregate datasets, models, and resources into a centralized and searchable repository.
Implement a multi-level validation and approval process to publish only verified datasets.
Define clear licensing terms for each dataset to ensure transparent usage and sharing rights.
Enable contributors to control dataset access, including open download, request-based access, or private use.
Provide a GPU-enabled sandbox environment to test datasets and train AI models directly on the platform.
The Solution
Unthinkable built AIKosh as a modular, scalable platform that unified datasets, AI tools, and computing resources. The initiative focused on AI data repository development to eliminate fragmented data sources, apply validation standards, and give practitioners easier access to high-quality datasets.
The team implemented structured data onboarding, licensing, controlled access, and GPU-enabled experimentation. By combining governance, collaboration, and infrastructure, the platform enabled faster model development, trusted datasets, and efficient national-scale AI innovation.
Key features developed were:
Unified Repository of Artefacts
- A unified hub for datasets, AI models, toolkits, and real-world AI use cases.
- Eliminates data silos across ministries, organizations, and research ecosystems.
- Enables researchers, startups, and enterprises to easily discover trusted AI resources.
- Advanced search by sector, organization, and tags improves discovery and knowledge sharing.

Role-Based Governance and Access Control
- Introduced role-based access: Explorer, Contributor, Organization Admin, and Platform Admin.
- Contributors upload datasets, AI models, and real-world use cases.
- Organization admins manage approvals and control access requests.
- Flexible sharing options: open, request-based, or private artefacts.

Integrated Sandboxing Environment
- GPU-enabled notebooks provide a ready sandbox for AI experimentation.
- Users can train and test models directly on the platform.
- Eliminates the need for costly external GPU infrastructure.
- Makes AI research faster, affordable, and more accessible.

Multi-Level Approval Workflow
- Every dataset and model undergoes a rigorous multi-step review process.
- Organization admins validate uploads before platform-level moderation.
- Final approval from the Ministry before publishing on the platform.
- Ensures authenticity, compliance, and trusted AI resources for users.

Flexible Data Ingestion
- Supports multiple dataset onboarding methods, including uploads, APIs, SFTP, and integrations.
- Integrates with platforms like GitHub, Kaggle, and Hugging Face.
- Simplifies contributions from government ministries and private organizations.
- API keys enable programmatic access for seamless developer workflows.

Data Quality And Validation
- Automated pipelines evaluate datasets for uniqueness, completeness, updates, and duplication.
- Ensures higher reliability and consistent dataset quality.
- Contributors receive star ratings based on data quality.
- Ratings motivate contributors to improve datasets continuously.

Licensing And Compliance
- Datasets are published under clearly defined licenses, from open-source to restricted terms.
- Ensures transparency in how data can be accessed, shared, or reused.
- Built-in anonymization tools protect sensitive information.
- Supports compliance with national data protection and ethical AI standards.

The Impact
AIKosh unified 10,000+ datasets and 200+ AI models from 60+ organizations across 20+ sectors, driving efficient AI data repository development. This centralization reduced discovery time, eliminated duplication, and enabled seamless access to trusted datasets.
The GPU-enabled notebook environment removed reliance on external infrastructure, lowered costs, and supported AI experimentation. Standardized access controls, transparent licensing, and automated quality scoring fostered secure cross-sector collaboration, building India’s scalable, future-ready AI ecosystem.
10000+
Datasets Aggregated
200+
AI Models Trained
60+
Contributing Organizations







