Batch Processing Concepts
Understand what "batch" means on a mainframe, how it differs from online/real-time processing, and the typical characteristics of a batch job: large volume, scheduled, unattended.
Introduction
You have used the word "batch" throughout this entire unit — batch job, batch processing, batch window — without pausing to define it precisely as its own concept. Now that you have hands-on familiarity with JCL, utilities, and how COBOL programs fit in, it is worth stepping back and looking at batch processing as a concept in its own right: what makes it "batch," what it is good at, and how it differs fundamentally from the real-time processing you will study later in the CICS unit.
What "Batch" Actually Means
Batch processing means running a program (or a sequence of programs) against a defined collection of data — a "batch" of records — from start to finish, without any user interactively waiting on or responding to individual results as they happen. A batch job is submitted, it runs, and it produces its results as a completed unit of work, typically minutes or hours later, for someone (or some downstream system) to pick up afterward.
The defining trait of batch processing is not speed or size by themselves — it is the absence of a person actively waiting on any single result while the work happens. A batch job can be small and fast or enormous and slow; either way, it runs to completion on its own, unattended.
Batch vs. Online/Real-Time
Contrasting batch against online (real-time) processing makes both concepts sharper. Online processing — which you will study in depth once you reach CICS later in this course — responds to individual requests as they arrive, one user or one event at a time, with an expectation of a near-instant reply. Batch processing instead works through a whole defined set of data as a single unit, with no individual response time expectation for any one record — only an expectation that the entire job finishes within its scheduled window.
| Aspect | Batch Processing | Online/Real-Time Processing |
|---|---|---|
| Trigger | Scheduled or submitted as a job | An individual user action or external event, at any time |
| Unit of work | An entire dataset or file, processed as a whole | One transaction at a time |
| Response expectation | The whole job finishes within its window (minutes to hours) | Sub-second response per transaction |
| Attendance | Unattended — nobody actively watches it run | A user or system is actively waiting on the result |
| Typical example | Nightly interest calculation across every account | A single customer checking their balance right now |
Characteristics of a Typical Batch Job
While no two batch jobs are identical, real-world mainframe batch work shares a recognizable set of traits worth naming explicitly.
- Large volume: batch jobs commonly process anywhere from thousands to many millions of records in a single run.
- Scheduled: batch jobs typically run at defined times — nightly, weekly, monthly, or end-of-quarter — rather than on demand from a user.
- Unattended: once submitted, a batch job runs to completion (or failure) without anyone actively monitoring or responding to it in real time.
- Sequential and repeatable: the same batch job, given the same input, is generally expected to produce the same output every time it runs.
- Dependent on prior steps or jobs: batch work is frequently organized into pipelines, where one job's output dataset is the next job's input.
A Familiar Example: End-of-Day Processing
A classic illustration is a bank's end-of-day batch cycle. During the business day, CICS handles online transactions in real time — deposits, withdrawals, transfers — each one applied against account records as it happens. Once the branch day closes, a large batch cycle runs: applying daily interest calculations across every account, generating statements, reconciling the day's transaction totals, and preparing files to feed regulatory or downstream reporting systems. Nobody sits and watches this happen record by record; the whole cycle is scheduled, monitored at a job level, and expected to complete before the next business day begins.
22:00 Batch window opens; CICS online activity has settled for the day22:05 JOB: EXTRACT - extract the day's posted transactions22:20 JOB: INTEREST - calculate and post daily interest across accounts22:50 JOB: STATEMENT - generate customer statement records23:30 JOB: RECONCILE - reconcile totals against the day's CICS activity23:50 JOB: FEED-REG - build files for downstream/regulatory reporting06:00 Batch window must be closed; branches reopen for businessClick Run to see what this code prints.
Why Batch Still Matters
It might be tempting to think of batch processing as the "old-fashioned" counterpart to modern real-time systems, but that framing misses why batch remains essential. Some kinds of work are fundamentally batch-shaped by nature: applying interest across every account at once, generating month-end reports, running large-scale data reconciliation, or processing enormous file transfers between institutions. These are not tasks anyone is waiting on moment to moment — they are tasks that need to be done completely, correctly, and on schedule, which is precisely what batch processing is built for.
Everything You Have Learned, Applied to Batch
It is worth explicitly noticing that this entire JCL and COBOL unit has been building the exact skill set batch processing depends on: JCL to describe and submit jobs, DD statements to connect programs to the datasets they process, PROCs to standardize recurring patterns, utilities like SORT and IDCAMS to handle common data operations, and COBOL programs to carry out the actual business logic against that data. Batch processing is not a separate topic from everything covered so far — it is the umbrella concept that everything you have learned in this unit exists to support.
Common Mistakes
- Assuming "batch" means "old" or "slow" — batch is a processing model suited to specific kinds of work, not a marker of outdated technology.
- Confusing large volume with the defining trait of batch — the real defining trait is the absence of a user actively waiting on individual results, regardless of the job's size.
- Assuming batch and online processing are rivals rather than partners — as the end-of-day example shows, they typically operate on the very same data at different times of day.
- Overlooking that batch jobs are frequently interdependent — treating each job as fully independent when it actually depends on a prior job's output can cause serious data or scheduling errors.
Best Practices
- When evaluating whether a task belongs in batch or online processing, ask whether anyone needs an individual, immediate response, or whether the whole task just needs to complete correctly within a window.
- Design batch jobs to be repeatable and predictable given the same input, since that predictability is what makes large-scale scheduling and troubleshooting manageable.
- Document dependencies between batch jobs clearly, since failures in shared, interdependent pipelines are far harder to diagnose without that context.
- Think of batch and online processing as complementary parts of one system, particularly when working with data that both touch, like core account records.
Frequently Asked Questions
Yes. Every JCL job you have submitted throughout this unit is, by definition, a batch job: it runs to completion unattended, without anyone waiting on individual results as it processes.
Yes. Volume is not the defining trait of batch processing — the absence of an interactive user waiting on individual results is. A batch job processing a handful of records is still batch, as long as it runs unattended, on its own.
Because batch work frequently needs exclusive or coordinated access to shared data (like account records also used by online CICS transactions during the day), and because downstream processes and business deadlines (statements, reports, regulatory feeds) depend on batch cycles finishing by specific times.
No. Certain categories of work — bulk calculations, large-scale reporting, reconciliation, and large file transfers — remain fundamentally batch-shaped, and mainframes continue to run enormous batch workloads alongside real-time CICS and web-facing systems.
Key Takeaways
- Batch processing means running a program against a whole set of data unattended, without any individual result being waited on interactively.
- Batch differs from online/real-time processing primarily in trigger, unit of work, and response expectations, not simply in volume or speed.
- Real-world batch jobs share common traits: large volume, scheduling, unattended execution, repeatability, and interdependency with other jobs.
- Batch and online processing typically coexist, often operating on the very same underlying data at different times of day.
- Everything covered earlier in this unit — JCL, DD statements, PROCs, utilities, COBOL — exists to support batch processing as a whole.
Summary
Batch processing is the umbrella concept behind everything you have practiced in this unit: unattended, scheduled work against large volumes of data, distinct from but complementary to real-time online processing. In the final lesson of this unit, you will look at how organizations actually coordinate large numbers of interdependent batch jobs using job schedulers, and at the concept of the batch window that constrains when all of that work has to fit.