Maycee Retail whitepaper

Maycee Retail Dataset

A synthetic Australian retail dataset for enterprise analytics, data engineering, and AI evaluation.

Prepared by SData | Version: v1.0 website edition

Executive summary

Modern organisations want to use AI to analyse business data, automate reporting, evaluate text-to-SQL systems, build dashboards, and test analytics agents. The practical challenge is that real enterprise data is hard to share safely. It can contain customer information, commercially sensitive trends, operational details, and governance constraints that make it unsuitable for public demos, training, prototyping, or external AI evaluation.

Maycee Retail Dataset was created by SData to provide a safer alternative: a realistic synthetic Australian retail analytics environment that can be used without exposing real customer or production data.

Maycee Retail includes a 13-table retail schema, daily partitioned Parquet data, synthetic customers, stores, products, promotions, transactions, item-level sales, returns, data quality scenarios, natural language enrichment, and starter question-to-SQL examples. It is designed for practical work across data engineering, analytics engineering, BI, machine learning, and AI evaluation.

The free tier is publicly available for 2017-2019 through Hugging Face under CC BY 4.0. Paid v1.0 release licences provide access to the current premium release snapshot for individual and team project use. Future benchmark subscription tiers will open only after their evaluator, verified task, answer-key, and model-report milestones are complete.

1. The problem

Enterprise AI and analytics systems need realistic data to be useful. A model can look impressive on a small tutorial dataset and still fail on real analytical work that involves joins, time periods, business definitions, missing values, customer segments, returns, promotions, and changing operational conditions.

  • Real production data is too sensitive for experimentation, training, vendor evaluation, or public demonstration.
  • Public sample datasets are often too small, too clean, or too generic to reflect enterprise complexity.
  • AI systems need repeatable evaluation data with known schema and expected behaviour.
  • Data teams need realistic pipeline and dashboard inputs before connecting to production systems.

2. What Maycee Retail provides

Maycee Retail simulates an Australian retail business environment. It is designed around analytical modelling patterns that data teams already use: dimensions, facts, relationships, daily partitions, business measures, and data quality checks.

13 retail schema tables.

10 dimensions and 3 fact tables.

Daily dt=YYYY-MM-DD partitions.

Free 2017-2019 public data.

Premium current-release data from 2020 onward.

Starter Q-to-SQL examples for AI and analytics-agent evaluation.

3. Design principles

Privacy-safe by design

Maycee Retail is synthetic. It does not contain real customer records, client data, production transactions, licence fulfilment details, or private operational material.

Enterprise-style structure

The schema is intentionally relational and analytical, supporting joins across customers, stores, products, promotions, dates, transactions, items, returns, and other business dimensions.

Analytics and AI readiness

The dataset is intended for dashboards, SQL practice, data quality checks, dbt-style modelling, text-to-SQL evaluation, RAG over schema documentation, and analytics-agent testing.

Australian retail context

The dataset is designed with Australian geography, fiscal-year use cases, retail seasonality, and event-style analytical patterns in mind.

4. Access model

Free tier

The free tier is public and available through Hugging Face under CC BY 4.0. It covers 2017-2019 and is designed for exploration, education, open examples, demos, and early technical evaluation.

Paid v1.0 release licences

For v1.0, SData sells release licences rather than subscriptions. Paid access is manually fulfilled by SData after payment, with access details and keys emailed within 1-2 business days after payment processing and fulfilment checks.

Tier Price Model Intended use
Individual Release Licence US$99 launch / US$149 later Perpetual current-release snapshot Personal projects, portfolio work, evaluation practice
Team Project Licence US$799 Up to 10 users, one internal project POCs, team demos, training, internal evaluation cycles

Individual and Team licences do not include support, future partitions, evaluator access, gold SQL answer keys, future benchmark tasks, model comparison reports, or hidden evaluation sets.

5. Use cases

Data engineering

Test ingestion pipelines, partition discovery, incremental loading, warehouse staging, transformation logic, orchestration patterns, and validation checks.

Analytics engineering

Build dbt-style models, star-schema marts, fiscal calendar handling, customer segmentation, product/category analysis, promotion effectiveness, returns analysis, and metric definitions.

Business intelligence

Power dashboards for sales performance, revenue trends, store comparisons, return rates, product categories, supplier performance, customer behaviour, and promotional uplift.

AI and text-to-SQL evaluation

Test whether models can generate useful SQL, interpret business questions, use the right joins, handle date logic, avoid unsupported claims, and explain results clearly.

6. Responsible AI value

Responsible AI requires evaluation before production use. For analytics agents, that means checking whether a system can generate SQL that runs, answer the intended business question, avoid making up missing metrics, handle ambiguity and dirty data, explain results clearly, and respect the limits of the available schema.

Maycee Retail provides a controlled, repeatable environment for this work. It lets teams test the behaviour of AI systems before connecting them to sensitive enterprise systems.

7. Roadmap

Maycee Retail v1.0 is live as a dataset and starter benchmark foundation. The larger benchmark layer will be released in stages.

v1.1 Benchmark MVP

Planned evaluator, 100 verified tasks, gold SQL, expected result hashes, dirty-data evaluation tasks, and first model comparison report.

v1.2 Benchmark Scale

Planned 400 verified tasks, harder and adversarial tasks, updated model comparison report, and full Professional benchmark subscription.

v2.0 Enterprise AI Benchmark

Planned RAG tasks, hallucination and insufficient-data tasks, dirty-data benchmark tasks, hidden evaluation set, model comparison dashboard, and enterprise evaluation workflows.

Enterprise Services remain a separate consulting engagement for organisations that need SData-led evaluation, custom synthetic datasets, private deployment, or advisory support.

8. Public and private boundaries

The public page and this whitepaper may describe Maycee Retail at a product and architecture level. They must not expose secrets, customer PII, customer-specific fulfilment records, licence manifests, paid customer access details, premium dataset files, private operational artefacts, gold SQL answer keys, hidden evaluation sets, evaluator internals, model reports, unapproved benchmark answers, client names, client logos, or private infrastructure details without explicit approval.

9. Why SData built Maycee

SData helps organisations build modern data, analytics, integration, AI, security, governance, and training capabilities. Maycee Retail is a productised demonstration of that delivery experience: a realistic dataset, a practical data model, a public learning asset, and a roadmap toward enterprise AI evaluation.

The goal is not to sell a simple data file. The goal is to give serious data and AI teams a reference environment they can use to build, test, demonstrate, and evaluate realistic workflows before working with production data.

10. Call to action

Start with the free public dataset on Hugging Face. Use the paid v1.0 release licences when you need current premium release data for individual or team project work. Talk to SData if your organisation needs AI evaluation, synthetic data design, private deployment, or modern data platform support.