LearnAI ToolsCareerPractice BuildsPlayContact
Mainframe SystemsBeginner~1.5 hours

Design a Dataset Layout

Plan sequential, PDS, and VSAM datasets for a simple student-records batch system.

Dataset DesignVSAMNaming Conventions

Overview

Before a single line of JCL runs a program, someone has to decide *what shape the data lives in* — and on z/OS that decision is much bigger than "a file" the way it is on Windows or Linux. A sequential dataset (PS) is read start to end like a flat file, with no way to jump to record 500 without reading the 499 before it. A partitioned dataset (PDS) is a directory of named "members," the mainframe's answer to a folder of small text files, and is where JCL and source code conventionally live. VSAM (Virtual Storage Access Method) datasets, especially the KSDS (Key-Sequenced Data Set) variant, behave much more like an indexed database table — records are stored and retrieved by a key, with `IDCAMS` doing for VSAM roughly what `CREATE TABLE` does for a relational database.

This project plans (not codes) the three datasets a minimal student-records batch system needs: a sequential file holding a day's raw enrollment transactions, a PDS holding the JCL and load modules that process them, and a VSAM KSDS holding the actual student master records, keyed by student ID so any one student can be looked up directly instead of scanned for. By the end you will have a naming convention, three dataset definitions, and a clear rule for which organization to reach for on a future project.

What You'll Build
  • A three-part high-level-qualifier naming convention shared by every dataset in the system.
  • A sequential (PS) dataset plan for daily raw enrollment transactions.
  • A PDS plan for JCL members and load modules, with directory block sizing.
  • An `IDCAMS DEFINE CLUSTER` plan for a VSAM KSDS student-master file, keyed by student ID.
  • A comparison table for choosing PS vs. PDS vs. VSAM KSDS on future projects.

Prerequisites

  • How to write and submit a basic JCL job (Project 1 in this course) — dataset allocation JCL builds directly on the JOB/EXEC/DD statements used there.
  • What DISP, SPACE, and DCB parameters describe on a DD statement.
  • The general idea of a key — a column or field that uniquely identifies one record, from any earlier database or programming background.
  • Familiarity with browsing a PDS in ISPF (option 3.4, dataset list).

Project Structure

The three datasets play distinct roles and are deliberately never the same organization type, because each one is optimized for how it is actually accessed: the transaction file is written once and read once, top to bottom, so sequential is both sufficient and the cheapest option; the JCL/program library is looked up by name on demand, which is exactly what a PDS directory is for; and the master file needs random lookup-by-key from an online or batch program, which only VSAM (or DB2) can do efficiently on z/OS.

DatasetOrganizationPurpose
STU.DAILY.TRANSSequential (PS)One day's raw enrollment transactions, processed top to bottom then discarded/archived
STU.PROD.JCLLIBPartitioned (PDS)JCL members for every batch job in the student-records system
STU.PROD.LOADLIBPartitioned (PDS)Compiled load modules the JCL in JCLLIB executes
STU.MASTER.KSDSVSAM KSDSOne record per student, keyed by student ID, updated in place

Step 1: Establish a Naming Convention

Every dataset name on z/OS is built from period-separated qualifiers, up to 44 characters total, and the leftmost one or two — the high-level qualifier (HLQ) — is what most shops use to group related data, apply storage-management rules, and control who is even allowed to see the name. Picking a convention before creating anything means every dataset's purpose is readable from its name alone, without opening it or checking documentation. The pattern below reads left to right as "project . environment/function . descriptive-name."

Qualifier PositionMeaningExample Values
1st (HLQ)Project/system identifierSTU (student-records system)
2ndEnvironment or functionDAILY, PROD, MASTER
3rdDescriptive dataset purposeTRANS, JCLLIB, LOADLIB, KSDS
Why this matters beyond readability

Storage groups, RACF dataset profiles (see Project 4), and automated cleanup jobs are almost always written against name patterns like `STU.PROD.*` — a sloppy or inconsistent naming convention means those rules either miss datasets they should cover or accidentally catch ones they shouldn't.

Step 2: Plan the Sequential Input Dataset

`STU.DAILY.TRANS` holds a full day's worth of enrollment change records — adds, drops, grade updates — one 80-byte record per transaction, in the order they were captured. It does not need to be cataloged permanently or organized for lookup, because the nightly batch cycle (Project 3) reads it once, start to finish, applies every transaction to the VSAM master, and then the file's job is done until tomorrow's replaces it. The allocation JCL below uses `IEFBR14`, a program that does nothing at all — it exists purely so a JCL step has something to "run" while the DD statement underneath it does the actual work of creating the dataset.

//ALLOCTR JOB (ACCT123),'J SMITH',CLASS=A,MSGCLASS=X
//STEP1 EXEC PGM=IEFBR14
//*
//* IEFBR14 does nothing but immediately return - it is the standard
//* "do-nothing" program used when a step's only purpose is allocating
//* or deleting a dataset via its DD statement, not running real logic.
//*
//TRANSDD DD DSN=STU.DAILY.TRANS,
// DISP=(NEW,CATLG,DELETE),
// SPACE=(TRK,(20,10),RLSE),
// UNIT=SYSDA,
// DCB=(RECFM=FB,LRECL=80,BLKSIZE=27920)
//*
//* One record per transaction, fixed 80-byte records (RECFM=FB).
//* 20 tracks primary / 10 secondary is sized for a few thousand
//* transactions/day; RLSE gives back any unused space once written.
//*

Step 3: Plan the PDS for JCL and Programs

A PDS adds one thing a plain sequential dataset does not have: a directory, so members can be stored and retrieved by name (`STU.PROD.JCLLIB(NIGHTRUN)`, for instance) instead of needing their own separate dataset name apiece. `DSORG=PO` (partitioned organization) is what actually makes it a PDS rather than a PS; the `SPACE` parameter's third sub-parameter, the directory blocks, is unique to partitioned datasets and has to be sized up front — 20 directory blocks holds roughly 200 small members, and unlike the data space itself, running out of directory blocks cannot be fixed by extending the dataset, only by reallocating it larger.

//JCLDD DD DSN=STU.PROD.JCLLIB,
// DISP=(NEW,CATLG,DELETE),
// DSORG=PO,
// SPACE=(TRK,(30,10,20)),
// UNIT=SYSDA,
// DCB=(RECFM=FB,LRECL=80,BLKSIZE=27920)
//*
//* DSORG=PO makes this a partitioned dataset (a PDS), not sequential.
//* SPACE=(TRK,(30,10,20)):
//* 30 tracks primary, 10 secondary (same as any dataset)
//* 20 <- directory blocks: space for the member-name directory itself,
//* separate from the data space and NOT extendable after the fact
//*
//LOADDD DD DSN=STU.PROD.LOADLIB,
// DISP=(NEW,CATLG,DELETE),
// DSORG=PO,
// SPACE=(TRK,(50,20,30)),
// UNIT=SYSDA,
// DCB=(RECFM=U,BLKSIZE=32760)
//*
//* Load modules use RECFM=U (undefined-length records) rather than FB,
//* since compiled program object code has no fixed record structure.
//*

Step 4: Plan the VSAM KSDS Master File

The student master needs something PS and PDS cannot give it: direct retrieval of one student's record by student ID, without reading everything before it. A VSAM KSDS stores records in key order and maintains an index over that key, so a program can `READ` by key and get the exact record in roughly constant time — the same access pattern an indexed database table gives you, built directly into the access method instead of a separate DBMS. Unlike PS/PDS allocation, a VSAM cluster is not defined with a plain `DD DISP=(NEW,...)` — it is defined with the `IDCAMS` utility's `DEFINE CLUSTER` command, run as a step's `SYSIN` control statements.

//DEFKSDS JOB (ACCT123),'J SMITH',CLASS=A,MSGCLASS=X
//STEP1 EXEC PGM=IDCAMS
//*
//* IDCAMS (Access Method Services) is the utility that defines,
//* deletes, and manages VSAM clusters - there is no plain-JCL DD
//* equivalent for "create a VSAM dataset" the way IEFBR14+DD works
//* for PS/PDS.
//*
//SYSPRINT DD SYSOUT=*
//SYSIN DD *
DEFINE CLUSTER (NAME(STU.MASTER.KSDS) -
INDEXED -
KEYS(9 0) -
RECORDSIZE(200 200) -
FREESPACE(10 10) -
VOLUMES(SYSDA1) -
TRACKS(100 50) ) -
DATA (NAME(STU.MASTER.KSDS.DATA)) -
INDEX(NAME(STU.MASTER.KSDS.INDEX))
/*
//*
//* INDEXED - defines a KSDS (as opposed to ESDS/RRDS variants)
//* KEYS(9 0) - the key is 9 bytes long, starting at offset 0
//* (student ID, e.g. '123456789')
//* RECORDSIZE(200 200) - every record is a fixed 200 bytes (avg, max)
//* FREESPACE(10 10) - leave 10% of each control interval/area free
//* so new records can insert in key order
//* without an immediate reorganization
//* TRACKS(100 50) - 100 tracks primary, 50 secondary allocation
//* Separate DATA and INDEX component names are given explicitly so
//* each can be referenced on its own in later JCL if ever needed.
//*
Why FREESPACE matters for a KSDS

Because records are stored in key order, inserting a brand-new student ID between two existing ones means physically making room in that exact spot. `FREESPACE(10 10)` reserves 10% of every control interval and control area for exactly this, so ordinary inserts do not immediately trigger a control-interval or control-area split — an expensive VSAM housekeeping operation that itself works best when it does not happen constantly.

Step 5: Choosing Between Dataset Types

With all three datasets planned, the underlying decision rule becomes clear: match the organization to the access pattern, not to habit. Sequential is the default for anything read start-to-end exactly once; partitioned is for anything looked up by a short human-assigned name; VSAM KSDS is for anything looked up by key from a program, especially when individual records need to be updated in place rather than rewritten as a whole file.

Access PatternBest FitWhy
Read once, top to bottomSequential (PS)No indexing overhead; cheapest to allocate and process
Look up a named member (JCL, source, load module)Partitioned (PDS/PDSE)Built-in directory gives name-based retrieval without a separate index dataset
Look up one record by key, from a program, possibly updating it in placeVSAM KSDSIndexed direct access by key; supports in-place update without rewriting the whole file
Rebuild the file from scratch overnight from a full extractSequential (PS)Simplicity wins when there is no need to preserve untouched records' physical position

Complete Plan

The full dataset layout, naming convention included, in the order these datasets would actually be provisioned before the nightly batch cycle in Project 3 can run against them.

Dataset NameOrganizationRecord FormatApprox. Size
STU.DAILY.TRANSSequentialFB, LRECL=8020 trk primary
STU.PROD.JCLLIBPartitionedFB, LRECL=8030 trk + 20 dir blocks
STU.PROD.LOADLIBPartitionedU, BLKSIZE=3276050 trk + 30 dir blocks
STU.MASTER.KSDSVSAM KSDSRECORDSIZE(200 200), KEYS(9 0)100 trk primary

Sample Run

Running the IEFBR14 allocation jobs and the IDCAMS DEFINE CLUSTER job in sequence against an empty system produces the following: three cataloged datasets and one VSAM cluster with its data and index components, all visible in an ISPF 3.4 dataset list afterward.

IDCAMS Execution Output

Click Run to see what this code prints.

Extend This Project

  • Add a fourth qualifier for test vs. production (`STU.TEST.DAILY.TRANS` vs `STU.PROD.DAILY.TRANS`) and update the naming-convention table.
  • Change `STU.PROD.LOADLIB` to a PDSE (`DSNTYPE=LIBRARY`) and explain in a comment why PDSEs avoid the "out of directory space" failure PDS libraries can hit.
  • Add an alternate index (`DEFINE AIX`) over the KSDS keyed by last name, for lookups where the student ID is not known.
  • Write the `IDCAMS DELETE` commands that would remove all four datasets cleanly, in the reverse order they were created.
  • Size `STU.DAILY.TRANS` for 50,000 transactions/day instead of a few thousand, and recompute the SPACE parameter.

Summary

You planned three dataset organizations for the same small system and matched each to how it is actually accessed rather than defaulting to one type everywhere: sequential for a file processed once top-to-bottom, partitioned for named JCL/program members, and VSAM KSDS for records that need direct lookup and in-place update by key. The naming convention and comparison table you built here are the same kind of decisions a mainframe systems programmer makes before any new batch subsystem is built, and the nightly batch cycle in the next project is designed to run directly against the datasets planned here.