GovTech / Data GovernanceShenzhen Smart City Group2026-08-22

SZSC: A Provincial Government Corpus Platform, 15 Industries x 12 Scenarios

As AI reaches government services, petabytes of raw government data have to become usable, high-quality corpora. SZSC built a provincial government corpus governance platform on MOI covering cleansing, annotation, quality inspection, classification and compliance review end to end.

Shenzhen Smart City Group

Shenzhen Smart City Group is a state-owned technology enterprise and one of China's leading smart city solution providers, operating across 160+ cities with 70+ subsidiaries, 15 certified high-tech enterprises and a 4,100-person technical workforce, 37% of whom hold master's or doctoral degrees.

First
Provincial corpus platform
15 x 12
Industries x scenarios
Petabyte
Raw data governed
End-to-end
Clean/annotate/QA/comply

The challenge

As AI transforms government services, SZSC faces a foundational challenge: turning petabytes of raw government data into high-quality, AI-ready corpus datasets that can power LLM-based applications across public safety, urban planning, transportation, healthcare, and citizen services. Government data is multi-modal (text, images, audio, video), scattered across dozens of agencies, subject to strict security and compliance requirements (classification, desensitization, audit trails), and must be processed through a rigorous pipeline-collection, cleaning, annotation, quality inspection, cataloging-before it can be used for model training or RAG applications.

The solution

SZSC built a Government Corpus Management Platform on OmniFabric-the first provincial-level government corpus governance platform-providing end-to-end corpus lifecycle management from raw data ingestion to AI-ready dataset delivery for downstream government AI applications.

The outcome

SZSC launched the first government-domain corpus governance platform in Guangdong Province, establishing the standard for how public-sector data is transformed into AI-ready assets.

Government data organized into 15 industry categories, 11 thematic classifications, and 12 application scenarios-giving AI developers structured, discoverable, compliance-cleared corpus at scale.

End-to-end security-data classification, dynamic desensitization, lineage tracking, sandbox isolation, and audit trails-ensures every corpus dataset meets government data protection standards.

Solution Architecture

Data sources
  • Raw government data (petabyte scale)
  • Archives and business system data
  • Heterogeneous documents and tables
  • Annotation and QA feedback
MatrixOne Intelligence
  • Cleansing and de-noising
  • Annotation and quality inspection
  • Classification: 15 industries x 12 scenarios
  • Compliance review and de-identification
  • Publication as AI-ready datasets
Business applications
  • Provincial government corpus library
  • Data supply for government AI apps
  • Compliance-traceable corpora
  • Cross-department reuse

Related solutions and products