Dummy Dataset Generation skill

Generate realistic dummy datasets for testing with customizable columns, constraints, and output formats (CSV, JSON, SQL, Python script).

by phuryn·MIT license·★ 26,557 Stars on the repo·GitHub ↗

Use now

Files of Dummy Dataset Generation

phuryn/main1 file shown
SKILL.md
Show the full text115 lines

Dummy Dataset Generation

Generate realistic dummy datasets for testing with customizable columns, constraints, and output formats (CSV, JSON, SQL, Python script). Creates executable scripts or direct data files for immediate use.

Use when: Creating test data, generating sample datasets, building realistic mock data for development, or populating test environments.

Arguments:

  • $PRODUCT: The product or system name
  • $DATASET_TYPE: Type of data (e.g., customer feedback, transactions, user profiles)
  • $ROWS: Number of rows to generate (default: 100)
  • $COLUMNS: Specific columns or fields to include
  • $FORMAT: Output format (CSV, JSON, SQL, Python script)
  • $CONSTRAINTS: Additional constraints or business rules

Step-by-Step Process

  1. Identify dataset type - Understand the data domain
  2. Define column specifications - Names, data types, and value ranges
  3. Determine row count - How many sample records needed
  4. Select output format - CSV, JSON, SQL INSERT, or Python script
  5. Apply realistic patterns - Ensure data looks authentic and valid
  6. Add business constraints - Respect business logic and relationships
  7. Generate or script data - Create executable output
  8. Validate output - Ensure data quality and completeness

Template: Python Script Output

import csv
import json
from datetime import datetime, timedelta
import random

# Configuration
ROWS = $ROWS
FILENAME = "$DATASET_TYPE.csv"

# Column definitions with realistic value generators
columns = {
    "id": "auto-increment",
    "name": "first_last_name",
    "email": "email",
    "created_at": "timestamp",
    # Add more columns...
}

def generate_dataset():
    """Generate realistic dummy dataset"""
    data = []
    for i in range(1, ROWS + 1):
        record = {
            "id": f"U{i:06d}",
            # Generate values based on column definitions
        }
        data.append(record)
    return data

def save_as_csv(data, filename):
    """Save dataset as CSV"""
    with open(filename, 'w', newline='') as f:
        writer = csv.DictWriter(f, fieldnames=data[0].keys())
        writer.writeheader()
        writer.writerows(data)

if __name__ == "__main__":
    dataset = generate_dataset()
    save_as_csv(dataset, FILENAME)
    print(f"Generated {len(dataset)} records in {FILENAME}")

Example Dataset Specification

Dataset Type: Customer Feedback

Columns:

  • feedback_id (auto-increment, U001, U002...)
  • customer_name (realistic names)
  • email (valid email format)
  • feedback_date (dates last 90 days)
  • rating (1-5 stars)
  • category (Bug, Feature Request, Complaint, Praise)
  • text (realistic feedback)
  • product (electronics, clothing, home)

Constraints:

  • Ratings skewed: 40% 5-star, 30% 4-star, 20% 3-star, 10% 1-2 star
  • Bug category only with ratings 1-3
  • Feature requests only with ratings 3-5
  • Email domains realistic (gmail, yahoo, company.com)

Output Deliverables

  • Ready-to-execute Python script OR direct data file
  • CSV file with proper headers and formatting
  • JSON file with valid structure and types
  • SQL INSERT statements for database population
  • Data validation and constraint compliance
  • Realistic, business-appropriate values
  • Documentation of data generation logic
  • Quick-start instructions for using the dataset

Output Formats

CSV: Flat tabular format, easy to import into spreadsheets and databases

JSON: Nested structure, ideal for APIs and NoSQL databases

SQL: INSERT statements, directly executable on relational databases

Python Script: Executable generator for custom or large datasets

1---
2name: dummy-dataset
3description: "Generate realistic dummy datasets for testing with customizable columns, constraints, and output formats (CSV, JSON, SQL, Python script). Use when creating test data, building mock datasets, or generating sample data for development and demos."
4---
5# Dummy Dataset Generation
6 
7Generate realistic dummy datasets for testing with customizable columns, constraints, and output formats (CSV, JSON, SQL, Python script). Creates executable scripts or direct data files for immediate use.
8 
9**Use when:** Creating test data, generating sample datasets, building realistic mock data for development, or populating test environments.
10 
11**Arguments:**
12- `$PRODUCT`: The product or system name
13- `$DATASET_TYPE`: Type of data (e.g., customer feedback, transactions, user profiles)
14- `$ROWS`: Number of rows to generate (default: 100)
15- `$COLUMNS`: Specific columns or fields to include
16- `$FORMAT`: Output format (CSV, JSON, SQL, Python script)
17- `$CONSTRAINTS`: Additional constraints or business rules
18 
19## Step-by-Step Process
20 
211. **Identify dataset type** - Understand the data domain
222. **Define column specifications** - Names, data types, and value ranges
233. **Determine row count** - How many sample records needed
244. **Select output format** - CSV, JSON, SQL INSERT, or Python script
255. **Apply realistic patterns** - Ensure data looks authentic and valid
266. **Add business constraints** - Respect business logic and relationships
277. **Generate or script data** - Create executable output
288. **Validate output** - Ensure data quality and completeness
29 
30## Template: Python Script Output
31 
32```python
33import csv
34import json
35from datetime import datetime, timedelta
36import random
37 
38# Configuration
39ROWS = $ROWS
40FILENAME = "$DATASET_TYPE.csv"
41 
42# Column definitions with realistic value generators
43columns = {
44 "id": "auto-increment",
45 "name": "first_last_name",
46 "email": "email",
47 "created_at": "timestamp",
48 # Add more columns...
49}
50 
51def generate_dataset():
52 """Generate realistic dummy dataset"""
53 data = []
54 for i in range(1, ROWS + 1):
55 record = {
56 "id": f"U{i:06d}",
57 # Generate values based on column definitions
58 }
59 data.append(record)
60 return data
61 
62def save_as_csv(data, filename):
63 """Save dataset as CSV"""
64 with open(filename, 'w', newline='') as f:
65 writer = csv.DictWriter(f, fieldnames=data[0].keys())
66 writer.writeheader()
67 writer.writerows(data)
68 
69if __name__ == "__main__":
70 dataset = generate_dataset()
71 save_as_csv(dataset, FILENAME)
72 print(f"Generated {len(dataset)} records in {FILENAME}")
73```
74 
75## Example Dataset Specification
76 
77**Dataset Type:** Customer Feedback
78 
79**Columns:**
80- feedback_id (auto-increment, U001, U002...)
81- customer_name (realistic names)
82- email (valid email format)
83- feedback_date (dates last 90 days)
84- rating (1-5 stars)
85- category (Bug, Feature Request, Complaint, Praise)
86- text (realistic feedback)
87- product (electronics, clothing, home)
88 
89**Constraints:**
90- Ratings skewed: 40% 5-star, 30% 4-star, 20% 3-star, 10% 1-2 star
91- Bug category only with ratings 1-3
92- Feature requests only with ratings 3-5
93- Email domains realistic (gmail, yahoo, company.com)
94 
95## Output Deliverables
96 
97- Ready-to-execute Python script OR direct data file
98- CSV file with proper headers and formatting
99- JSON file with valid structure and types
100- SQL INSERT statements for database population
101- Data validation and constraint compliance
102- Realistic, business-appropriate values
103- Documentation of data generation logic
104- Quick-start instructions for using the dataset
105 
106## Output Formats
107 
108**CSV:** Flat tabular format, easy to import into spreadsheets and databases
109 
110**JSON:** Nested structure, ideal for APIs and NoSQL databases
111 
112**SQL:** INSERT statements, directly executable on relational databases
113 
114**Python Script:** Executable generator for custom or large datasets
115 

Discussion

Alternatives

AI engineerAct as an expert AI engineer specializing in practical machine learning implementation and AI integration for production applications, ensuring efficient and robust AI solutions.Data & AI · CC0-1.0OneKGPd: Individual-Level Queries over the 1000 Genomes ProjectQuery the 1000 Genomes Project dataset (3,202 whole-genome-sequenced individuals, GRCh38) at the level of individual participants. Use when a question is about individuals or variants in the 1000 Genomes Project cohort: which individuals carry variants matching specific criteria in a gene or region, which individuals are homozygous-reference at a position, which variants exist in the dataset or carried by specified individuals in a gene or region, the relatedness between two specified individuals. Variants are returned with 1000 Genomes allele frequencies (AF), gnomAD v4.1 exome and genome AF, AlphaMissense score, and HGVSp annotations.Science · MITPyMC Bayesian ModelingBayesian modeling with PyMC. Build hierarchical models, MCMC (NUTS), variational inference, LOO/WAIC comparison, posterior checks, for probabilistic programming and inference.Science · MITStatsmodels: Statistical Modeling and EconometricsStatistical models library for Python. Use when you need specific model classes (OLS, GLM, mixed models, ARIMA) with detailed diagnostics, residuals, and inference. Best for econometrics, time series, rigorous inference with coefficient tables. For guided statistical test selection with APA reporting use statistical-analysis.Science · MIT