South Korea’s government opened the vault on one of the largest publicly available AI training datasets ever released by a national government on August 27, 2026. The Ministry of Science and ICT (MSIT), working with the National Information Society Agency (NIA), announced that 35.44 million data records totaling roughly 1.56 trillion tokens are now downloadable, free, through the state-run AI Hub portal. The datasets come from five companies competing in Korea’s “sovereign AI foundation model” contest, and the move is the clearest signal yet that Seoul intends to buy its way into the frontier AI race with public infrastructure rather than wait on Silicon Valley.
The release, first reported by ChosunBiz and confirmed by the Seoul Economic Daily’s English edition, lands just two weeks after a state audit exposed serious quality problems in South Korea’s earlier AI data spending. That timing matters. Korea is trying to prove that its second attempt at public AI data is cleaner, bigger and more useful than its first, and the whole industry is watching whether 1.56 trillion tokens of government-backed data can actually move the needle for companies like Naver and LG that are chasing GPT-5.6 and Gemini-class capability.
What South Korea actually released on August 27
MSIT and NIA said they are publishing 29 distinct dataset types built by the five teams that cleared the first-stage evaluation of Korea’s sovereign AI foundation model project, an internal government contest sometimes shorthanded in Korean coverage as “독파모” (dokpamo, short for “독자 AI 파운데이션 모델,” or independent AI foundation model). Those five teams are Naver Cloud, Upstage, SK Telecom, NC AI and LG AI Research, according to Sedaily’s reporting.
The headline numbers, confirmed across ChosunBiz, Sedaily and the specialist legal-and-policy wire MLex, are 35.44 million records and approximately 1.56 trillion tokens. ChosunBiz reports the 2025 budget for building this specific batch of data at 15 billion won, or roughly $10.8 million at current exchange rates. That is a comparatively modest sum next to the scale of the release, which is the point Korean officials are making: this data was collected as a byproduct of a competitive process already underway, not commissioned from scratch.
MLex reports the dataset is large enough to pretrain foundation models in the 70-billion to 80-billion parameter range, and that it spans more than plain text. The package includes pretraining corpora, multimodal data covering video and audio, and red-teaming data specifically built for safety evaluation. That last category matters for a government trying to avoid the reputational hit of shipping a sovereign model that fails basic safety benchmarks.
Who gets access, and who doesn’t
Access sits behind a straightforward gate. Any domestic company, researcher, or student can search, download and use the released data free of charge, filed under the “Sovereign AI Model Data” category on AI Hub, the government’s public data platform run jointly by MSIT and NIA. That is a meaningfully open policy compared to typical government data releases, which often require formal applications, NDAs, or per-project approval.
Not everything is unrestricted, though. Some of the more sensitive datasets in the package must be requested separately and used only inside a secured environment the government calls “Safety Zone,” rather than downloaded directly. This two-tier system, general-access data on one side and gated data behind Safety Zone on the other, is Korea’s attempt to balance openness with the privacy and security risks baked into large multimodal training corpora, particularly the video and audio components MLex flagged.
The ministry also said that data built during the second stage of the sovereign AI foundation model evaluation will be released later, once it clears quality verification. That is a notable caveat given what happened with the last round of government AI data, which is worth unpacking before assessing whether this release will actually work.
The audit that came two weeks earlier
This release did not happen in a vacuum. On August 12, 2026, a state audit revealed that South Korea had already spent about 1.6328 trillion won, roughly $1.13 billion, building 908 kinds of AI training data between 2017 and 2024, all published through the same AI Hub platform. Korea JoongAng Daily reported that the audit found duplicate and unusable data scattered across that earlier program, a finding independently confirmed by UPI and The Chosun Daily’s English edition, which ran the story under the headline “Government’s 1.6T AI Data Mostly Unusable.”
The audit’s core criticism was coordination, not intent. Multiple agencies had independently commissioned overlapping datasets without checking what already existed on AI Hub, producing redundant records and, in some cases, data that failed basic quality checks needed for model training. That is a billion-dollar cautionary tale sitting directly behind the August 27 announcement, and it explains why MSIT built explicit quality-verification language into this new release rather than simply dumping the files online.
Whether the new 1.56-trillion-token package avoids the same duplication and quality problems is not yet independently verified. No outlet covering the August 27 release has published an outside audit of the new data’s quality, only the ministry’s own characterization of it as vetted, contest-tested material from five commercial AI labs rather than the more scattershot multi-agency sourcing that produced the earlier 908-dataset pile.
Why South Korea is doing this: the sovereign AI push
The sovereign AI foundation model project is Korea’s answer to a strategic worry shared by a growing list of mid-sized economies: if you don’t build your own frontier-class large language model, you end up permanently renting intelligence from OpenAI, Google, Anthropic or a Chinese lab, with all the data-sovereignty and pricing exposure that implies. Japan, France, the UAE and India have all launched comparable national-champion AI programs over the past two years, each betting that domestic language, culture and regulatory needs justify a homegrown model even when it can’t match GPT-5.6 or Gemini 3 on raw benchmark scores.
Korea’s version runs as a staged competition. Five teams, Naver Cloud, Upstage, SK Telecom, NC AI and LG AI Research, cleared the first-stage evaluation, and the training data those teams built along the way is now the public asset being released. That structure means the government effectively crowdsourced its national AI dataset from the private sector’s best-funded labs, then repackaged the output as public infrastructure. It is a clever way to get five well-capitalized companies to build a shared national resource without directly commissioning it, and it sidesteps some of the coordination failures that plagued the earlier 908-dataset program.
How 1.56 trillion tokens compares to the world’s biggest open datasets
Context matters here, and the honest answer is that 1.56 trillion tokens, while large for a single government release, is not close to the biggest open text corpus in existence. Hugging Face’s FineWeb dataset, maintained by the open-source AI community, contains more than 18.5 trillion tokens on its own, according to the dataset’s official documentation. FineWeb-Edu, a filtered education-focused subset, holds about 1.3 trillion tokens. Older but still widely used corpora like RedPajama (roughly 1.02 trillion tokens) and The Pile (about 286 billion tokens) sit in a similar or smaller range than Korea’s new release.
What makes the Korean release distinctive isn’t raw scale against FineWeb, it’s the fact that no government appears to have matched it as a single, centralized, publicly released text-and-multimodal corpus. The U.S. National Science Foundation has said explicitly that it runs no dedicated programs supporting “operational national-scale data systems,” leaning instead on the National AI Research Resource pilot, which grants access to third-party datasets like AI2’s Dolma rather than building a federal one. The European Union’s public AI data infrastructure centers on data spaces and curated vocabularies rather than one flagship training corpus, and available reporting on China points to a centralized dataset management platform under construction rather than a single dataset already published at comparable scale. On that specific axis, government-built-and-released rather than merely government-funded, Korea’s move looks closer to a first than a follower.
Data table: South Korea’s AI Hub release by the numbers
| Metric | Figure | Source |
|---|---|---|
| Total tokens released | ~1.56 trillion | ChosunBiz, Sedaily, MLex |
| Total records/items | 35.44 million | ChosunBiz, Sedaily |
| Dataset types included | 29 | ChosunBiz, Sedaily |
| Contributing teams | 5 (Naver Cloud, Upstage, SK Telecom, NC AI, LG AI Research) | Sedaily |
| 2025 build budget for this batch | 15 billion won (~$10.8M) | ChosunBiz |
| Estimated model scale supported | ~70B–80B parameters | MLex |
| Release date | August 27, 2026 | ChosunBiz, Sedaily |
| Access model | Free for domestic firms, researchers, students via AI Hub; sensitive sets gated in “Safety Zone” | Sedaily |
| Prior program spend (2017–2024) | 1.6328 trillion won (~$1.13B) for 908 dataset types | Korea JoongAng Daily, UPI, Chosun |
Data table: how it stacks up against other public training corpora
| Dataset | Approx. token count | Source/maintainer |
|---|---|---|
| Hugging Face FineWeb | 18.5 trillion | Hugging Face (open community) |
| South Korea AI Hub (Aug 2026 release) | 1.56 trillion | MSIT / NIA (government) |
| Hugging Face FineWeb-Edu | 1.3 trillion | Hugging Face |
| RedPajama | ~1.02 trillion | Together AI / community |
| The Pile | ~286 billion | EleutherAI |
The comparison table underscores a point worth repeating: this isn’t the largest dataset on the planet, it’s the largest one a national government has directly assembled and shipped as a single public release, aimed squarely at closing the gap between Korean labs and the US and Chinese frontier.
What officials are saying
The Ministry of Science and ICT and the National Information Society Agency framed the release plainly in their joint statement: “The Ministry of Science and ICT and the National Information Society Agency said on the 27th that they will release 29 datasets on AI Hub, built by the five elite teams that participated in the first-stage evaluation of the sovereign AI foundation model project — NAVER Cloud, Upstage, SK Telecom, NC AI and LG AI Research,” according to Sedaily’s report.
On who can use it, the same joint statement was equally direct: “Any domestic company, researcher or student can search, download and use the released data free of charge under the ‘Sovereign AI Model Data’ category on AI Hub,” per Sedaily.
And on what comes next, MSIT signaled this is a first tranche, not the whole program: “The ministry said data newly built during the second-stage evaluation will also be released after quality verification,” according to the same Sedaily reporting. That “after quality verification” clause reads directly as a response to the August 12 audit criticism, and it’s the clearest evidence that the ministry is trying to avoid a repeat of the duplication problems found in the earlier 908-dataset program.
Market and industry impact
None of the outlets covering the August 27 release reported a stock-market reaction, and that’s not surprising given the data went to five named contest teams rather than triggering a surprise corporate announcement. The more meaningful impact is structural. Naver Cloud, Upstage, SK Telecom, NC AI and LG AI Research are the five companies Korea has effectively designated as its national AI champions through the sovereign foundation model contest, and this data release lowers the training-data cost curve for every smaller Korean AI startup and academic lab that wants to build on top of what those five teams already assembled.
That’s a meaningful subsidy in an industry where data acquisition and cleaning routinely eats a large share of a startup’s early compute budget. A 15-billion-won public investment that produces 1.56 trillion tokens of usable, multimodal, safety-annotated training data is, on paper, an efficient way for the government to seed an entire domestic ecosystem rather than fund one company’s model directly. Whether Korean AI startups outside the five contest teams actually adopt this data at scale, versus continuing to rely on scraped web corpora and licensed content, is the open question the next two quarters should answer.
The quality question nobody has answered yet
The elephant in the room is quality, and it is the single biggest risk to this program’s credibility. The August 12 audit didn’t just flag duplication in the prior 908-dataset program, it found data that was “difficult or impossible to use for model training,” according to reporting synthesized from the Chosun, Korea JoongAng Daily and MLex coverage of that audit. If a meaningful fraction of the new 1.56-trillion-token package inherits similar problems, and no outside party has yet audited it, the headline token count could prove misleading in practice, since garbage tokens don’t train useful models any better than no tokens at all.
There’s a structural reason for cautious optimism, though. This new data came out of an active, judged competition among five well-resourced commercial labs, each with a direct incentive to build data that actually improves their own model’s contest performance. That’s a fundamentally different incentive structure than the prior program, where multiple government agencies independently commissioned data with no shared quality bar and no competitive pressure to make it good. Incentive alignment doesn’t guarantee clean data, but it’s a more defensible starting point than what produced the $1.13 billion pile the auditors picked apart.
Privacy, copyright and the Safety Zone compromise
Large multimodal datasets built from real-world video and audio inevitably raise privacy questions, and Korea’s two-tier access system is a direct acknowledgment of that risk. General datasets are free and open on AI Hub; anything sensitive enough to carry privacy or security exposure is walled off in the Safety Zone environment, requiring separate request and controlled, in-environment use rather than a direct download. That’s a more conservative posture than a fully open release, and it will likely slow adoption of the sensitive subsets, but it’s also the kind of guardrail regulators in the US and EU have been pushing AI labs toward voluntarily, mostly without success.
Copyright is the quieter risk. None of the current reporting details exactly how the underlying source material for these 29 datasets was licensed or cleared, which matters given how aggressively rights holders have pursued AI training-data lawsuits in the US and Europe over the past two years. If any of the 35.44 million records trace back to copyrighted Korean media, publishing, or broadcast content without clear licensing, that’s a legal exposure that could surface well after the celebratory headlines fade.
Historical context: Korea’s AI data ambitions since 2017
South Korea has been building public AI training data infrastructure for nearly a decade. The AI Hub platform, run by MSIT and NIA, has been the delivery mechanism for that entire nine-year effort, and the August 12 audit’s finding of 908 dataset types built between 2017 and 2024 shows just how much has flowed through it already. That’s an average of well over 100 new dataset types published per year, at an average cost of roughly $1.24 million each, a pace that outstrips almost any other country’s public AI data program in sheer volume of distinct efforts, even if quality control lagged behind.
The sovereign AI foundation model contest that produced this new 1.56-trillion-token release represents a shift in method rather than ambition. Instead of commissioning data directly through scattered agency contracts, as the earlier 908-dataset program did, the government is now harvesting data as a byproduct of a competitive process it funds and judges. If that structural change actually fixes the duplication and usability problems the auditors found, it could become a template other AI Hub programs adopt going forward.
What this means for Naver, Upstage, SK Telecom, NC AI and LG AI Research
For the five contest teams, this release is a mixed blessing. On one hand, publishing your training data publicly hands smaller Korean competitors and international startups a leg up on data assembly cost, potentially eroding one of the advantages that comes with being a well-capitalized incumbent. On the other hand, the government is effectively subsidizing the entire domestic AI ecosystem’s data costs using material these five companies already built for their own contest submissions, which lowers barriers for future partners, downstream applications and academic collaborators who might otherwise never work with Naver Cloud or LG AI Research directly.
It’s also a soft signal about where the government’s sovereign AI bet is concentrated. Naver Cloud and LG AI Research already run publicly known foundation model programs (HyperCLOVA X and EXAONE, respectively), while SK Telecom, NC AI and Upstage bring telecom-scale infrastructure, gaming-industry compute experience, and enterprise LLM tooling respectively. Having all five clear the first-stage evaluation suggests Korea is deliberately keeping multiple horses in the sovereign-AI race rather than picking a single national champion this early, a hedge that costs more in the short term but reduces the risk of the entire program collapsing if any one company’s model underperforms.
Predictions: where this goes next
- Second-stage data will land within one to two quarters. MSIT explicitly promised additional data after second-stage evaluation and quality verification, which likely puts the next tranche somewhere between late 2026 and early 2027.
- Independent quality audits of the new 1.56T-token package are coming. Given the fallout from the August 12 audit, opposition lawmakers or the Board of Audit and Inspection are likely to review this new release within the next year to confirm it avoids the same duplication problems.
- At least one of the five contest teams will ship a public benchmark result citing this dataset. Expect Naver Cloud or LG AI Research to reference AI Hub-sourced training data in technical reports for their next HyperCLOVA X or EXAONE-class model release.
- Other mid-sized AI nations will study Korea’s contest-based data model. Countries like Japan, the UAE, and Singapore running their own sovereign-AI programs are likely to evaluate whether harvesting training data from a judged competition, rather than direct commissioning, produces better results than their current approach.
- Copyright scrutiny is likely within the next 12 months. As AI training-data litigation intensifies globally, expect Korean media or publishing groups to at least request clarity on how the 35.44 million records were sourced and licensed.
The bigger picture: does public AI data actually work?
South Korea’s bet is that publicly funded, openly released training data can meaningfully accelerate a national AI industry without the government having to build or own a frontier model itself. It’s a lower-risk strategy than, say, the UAE’s direct funding of Falcon or France’s backing of Mistral, but it’s also a slower one, success depends on private companies actually choosing to build on top of what the state hands them rather than defaulting to proprietary or scraped data they control end to end.
The August 12 audit is the cautionary half of this story, proof that good intentions and a nine-figure-plus budget don’t automatically produce usable AI infrastructure. The August 27 release is the optimistic half, a bet that competition-sourced data, built by companies with skin in the game, avoids the coordination failures that plagued the earlier effort. Whether Korea actually closes the gap with GPT-5.6, Gemini and Claude-class models will depend less on the token count in this week’s headlines and more on what happens to data quality once outside auditors get a look at it.
Frequently asked questions
How many tokens is South Korea’s new AI training dataset?
Roughly 1.56 trillion tokens across 35.44 million records, spanning 29 dataset types, according to the Ministry of Science and ICT and National Information Society Agency’s August 27, 2026 announcement.
Who can access South Korea’s AI Hub training data?
Any domestic company, researcher, or student can search, download and use the general datasets free of charge through the “Sovereign AI Model Data” category on AI Hub. Some sensitive datasets require separate approval and must be used within a secured “Safety Zone” environment rather than downloaded directly.
Which companies built the data being released?
Five companies competing in Korea’s sovereign AI foundation model contest: Naver Cloud, Upstage, SK Telecom, NC AI and LG AI Research. All five cleared the contest’s first-stage evaluation, and the training data they assembled along the way is now the public release.
How much did this data cost to build?
ChosunBiz reports the 2025 construction budget for this specific dataset batch at 15 billion won, roughly $10.8 million. That’s separate from the 1.6328 trillion won (about $1.13 billion) South Korea spent on AI training data broadly between 2017 and 2024, per the August 2026 state audit.
Is this the largest public AI training dataset in the world?
No. Hugging Face’s community-maintained FineWeb dataset holds more than 18.5 trillion tokens, over ten times larger. What makes Korea’s release notable is that it appears to be the largest single dataset a national government has directly assembled and released as public infrastructure, rather than the largest dataset overall.
Was South Korea’s earlier AI data criticized for quality problems?
Yes. An August 12, 2026 state audit found that South Korea’s prior AI Hub program, which built 908 kinds of training data at a cost of about $1.13 billion between 2017 and 2024, contained significant duplication and datasets that were difficult or impossible to use for model training, according to Korea JoongAng Daily, UPI and The Chosun Daily.
What model sizes can this data train?
MLex reports the dataset is large enough to pretrain foundation models in the 70-billion to 80-billion parameter range, and includes multimodal data covering video and audio, plus red-teaming data built for safety evaluation.
Will more data be released later?
Yes. MSIT said data built during the sovereign AI foundation model contest’s second-stage evaluation will also be released, once it passes quality verification, though no specific release date has been announced.
Related Coverage
- DeepSeek V4 vs R1 vs V3.2: Peak Prices Surge 355% [2026]
- Nvidia Rubin vs AMD Helios vs Microsoft Maia 300: AI Chip Race Hits $100B [2026]
- Cerebras CS-4 Claims 30x Faster AI Inference Than GPUs [2026]
- GPT-5.6 vs DeepSeek V4 Pro 0813: 714x Cheaper Input [2026]
- OpenAI Report on Hugging Face AI Agent Hack: 4 Services Hit [2026]
- Claude Opus 5 vs GPT-5.6 vs DeepSeek V4-Pro: $22 Gap [2026]