Maycee product family

Maycee Retail Dataset

A synthetic Australian retail dataset for data engineering, analytics, dimensional data modelling, BI development, text-to-SQL testing, and AI agent evaluation.

Maycee Retail gives technical teams a realistic, privacy-safe retail analytics environment without exposing customer or commercial production data. It combines a 13-table retail schema, daily Parquet partitions, documented metadata, data quality scenarios, natural language enrichment, and starter question-to-SQL examples.

13

schema tables

10

years unified (3 free + 7 premium)

1,095

free daily partitions

22

starter Q-to-SQL pairs

Why it exists

Modern data and AI teams need realistic data environments, not toy datasets.

AI teams, analytics engineers, BI developers, consultants, vendors, and educators often need realistic data for demonstrations, training, prototyping, and evaluation. Real enterprise data is difficult to share because it can contain private customer details, sensitive operational patterns, or commercial information.

Maycee Retail is designed to bridge that gap. It is synthetic, Australian, retail-oriented, and structured for enterprise-style analytics work. Teams can use it to practise data modelling, test pipelines, build dashboards, evaluate text-to-SQL behaviour, and demonstrate analytics workflows without using sensitive production data.

What is included

Built for practical analytics and AI evaluation workflows.

Retail star schema

Transactions, items, returns, and 10 dimensions covering dates, regions, districts, suppliers, brands, categories, customers, stores, products, and promotions.

Partitioned Parquet

Daily dt=YYYY-MM-DD partitions designed for lakehouse, warehouse, and incremental pipeline workflows.

AI-ready context

Natural language fields, documented schema and metadata, starter Q-to-SQL pairs, and roadmap benchmark assets for evaluating analytics agents.

Dirty-data scenarios

Intentional data quality scenarios for testing validation logic, pipeline resilience, and model behaviour under imperfect data.

Dashboard and modelling use cases

Suitable for BI dashboards, dimensional modelling, dbt-style transformations, SQL notebooks, training workshops, and data platform demonstrations.

Privacy-safe synthetic data

No real customer records, production transactions, or client-specific delivery data are included.

Free tier

Start with the free public dataset.

The free tier covers 2017-2019 and is available publicly through Hugging Face. It is the fastest way to explore the schema, inspect the Parquet files, and test basic analytics or AI workflows before choosing a paid release licence.

Access free dataset on Hugging Face
Coverage2017-2019 public dataset
Partitions1,095 daily partitions
Schema13-table retail model
LicenceCC BY 4.0, no account required

Paid v1.0 release licences

Buy the current Maycee Retail v1.0 release.

Paid v1.0 licences provide perpetual access to Maycee Retail as it exists at the time of purchase. No subscription tiers are sold at v1.0.

Public access

Free

US$0

2017-2019 public dataset on S3 / Hugging Face

Developers, students, lead generation

Public 2017-2019 dataset, basic benchmark tasks, documentation, no support.

One-off project licence

Team Project Licence

US$799

Up to 10 users, one internal project, perpetual access to current release, no updates

Small teams, POC and evaluation projects

Up to 10 users for one internal project. Perpetual access to Maycee Retail as it exists at time of purchase, all partitions up to that date, current Q->SQL pairs, whitepaper and documentation. No future partitions, evaluator access, model reports, or support.

Fulfilment note

Paid access is manually fulfilled by SData after payment. SData will email access details and keys within 1-2 business days. Customers do not need to log in to the SData website for v1.0 dataset access.

Licence boundary

Paid licences do not permit redistribution, resale, public re-upload, or sharing outside the licensed individual or team/project boundary. Individual and Team licences do not include future partitions, future benchmark tasks, evaluator access, gold SQL answer keys, model comparison reports, or support.

Who it is for

Useful for teams that need realistic data without production risk.

Data engineers testing ingestion, partitioning, validation, and warehouse workflows.

Analytics engineers building dbt-style models and semantic layers.

BI developers building dashboards and executive reporting examples.

AI engineers evaluating text-to-SQL, RAG, and enterprise analytics agents.

Consultants demonstrating modern data and AI platform patterns.

Training providers teaching realistic analytics and AI workflows.

Benchmark roadmap

The dataset is live. The benchmark layer is the roadmap.

Maycee Retail v1.0 includes starter Q-to-SQL examples and AI-ready documentation. The full benchmark product is staged for later releases, with detail provided in the attached whitepaper roadmap. SData will not describe evaluator, gold SQL, model reports, hidden evaluation sets, or live leaderboard features as available until those assets are complete.

Whitepaper

Read the Maycee Retail whitepaper.

The whitepaper explains the dataset design, schema, use cases, responsible AI rationale, licensing model, and benchmark roadmap in more detail.

Read whitepaper

Public boundary

The public page and whitepaper describe Maycee Retail at a product and architecture level. They do not expose premium data files, licence manifests, customer records, secrets, private infrastructure details, gold SQL, hidden evals, evaluator internals, or model reports.

Build and test with realistic synthetic retail data.

Start with the free public dataset on Hugging Face, buy a v1.0 release licence for premium project use, or talk to SData about enterprise AI evaluation and synthetic data services.