Data architecture · Planning guide
Design the platform around data products and consumers.
A modern data platform connects ingestion, storage, transformation, orchestration, governance, observability, analytical models, and serving layers into a system that can support reporting, applications, analytics, and AI.
Who this is for
Data architects, domain owners and platform teams designing shared data products and consumer interfaces.
What to leave with
A layered platform design with data ownership, quality contracts and serving requirements.
Workflow design
Reference approach: Modern data platform architecture
Use this sequence to identify interfaces, review points and evidence. Adapt the stages to your systems; it is a planning reference, not a client result.
- 01
Capture source changes
Choose ingestion patterns based on source behavior, delivery windows and replay requirements.
Output
Traceable source data with freshness context
- 02
Store and transform
Separate raw inputs from curated models and preserve lineage through tested transformations.
Output
Versioned, quality-checked datasets
- 03
Publish data products
Expose stable definitions and access-controlled interfaces for analytical or application consumers.
Output
Consumer-ready data contracts
- 04
Operate across layers
Monitor freshness, quality and failures with named ownership and an actionable recovery path.
Output
An observable platform operating model
Controls across the workflow
- Named source and workflow owners
- Reviewable acceptance evidence
- Explicit access and operating boundaries
- Recorded exceptions and next actions
Decisions that shape the scope
- Are data domains and owners explicit?
- Assign responsibility for source meaning, transformation logic and the published data contract.
- Are quality failures actionable?
- Identify the signals, owners and recovery process for failed tests, stale sources and broken consumer interfaces.
A modern data platform is an operating architecture, not a collection of tools
Cloud warehouses, lakehouses, transformation frameworks, orchestration tools, catalogues, and BI platforms are useful components, but the architecture should begin with data domains and consumers. Sources need clear ingestion patterns, datasets need ownership, transformations need testing, and downstream users need stable interfaces and trustworthy definitions.
Decisions to work through
01
Design around data domains
Organise datasets, ownership, quality, and serving responsibilities around meaningful business or operational domains.
02
Separate platform layers
Make ingestion, storage, transformation, governance, serving, and consumption responsibilities explicit rather than hiding them inside individual tools.
03
Treat quality as infrastructure
Testing, lineage, freshness, observability, contracts, and ownership should be part of normal platform operation.
04
Serve multiple consumers
Design curated outputs for BI, operational applications, analysts, data science, machine learning, and AI systems according to their different needs.
Review before you proceed
Use this checklist to structure the discussion. Ticking an item records your review here; it does not certify readiness. Your selections reset when you reload.
0 of 4 reviewed
Comparison
Core layers of a modern data platform architecture
| Area | What to evaluate | Why it matters |
|---|---|---|
| Sources | Operational databases, SaaS systems, APIs, files, events, applications, and external data. | Source ownership, extraction limits, schema change, security, and data authority affect everything downstream. |
| Ingestion | Batch loads, CDC, streaming, API extraction, file ingestion, and synchronisation. | Pipelines need reliability, replay, observability, and clear freshness expectations. |
| Storage | Warehouse, lake, lakehouse, object storage, operational stores, and specialised databases. | Storage choices should reflect workload, governance, access, performance, and lifecycle requirements. |
| Transformation | Cleaning, standardisation, joins, business rules, models, tests, and reusable transformations. | This is where raw data becomes dependable analytical and application-ready information. |
| Orchestration | Scheduling, dependencies, retries, event triggers, workflow state, and pipeline coordination. | Reliable data products require predictable execution across many dependent processes. |
| Governance | Ownership, catalogue, lineage, classifications, permissions, policies, and retention. | Without governance, platform scale can increase confusion faster than value. |
| Observability | Freshness, volume, quality, schema change, failures, lineage impact, and operational alerts. | Data incidents need to be detected before business users discover them. |
| Serving | Semantic models, BI, APIs, reverse ETL, applications, ML features, search, and AI context. | The platform exists to deliver useful and reliable data to downstream consumers. |
Scroll horizontally to view the full comparison on smaller screens.
Frequently asked questions
Modern data platform architecture: questions and answers
01What is a modern data platform?
What is a modern data platform?
A modern data platform is an architecture for collecting, storing, transforming, governing, observing, and serving organisational data across analytics, operational applications, data science, machine learning, and AI workloads.
02Does a modern data platform require a lakehouse?
Does a modern data platform require a lakehouse?
No. Lakehouse architecture is one option. The appropriate design may use warehouses, lakes, lakehouses, operational databases, specialised stores, or combinations according to workload requirements.
03What are the main layers of a data platform?
What are the main layers of a data platform?
Common layers include sources, ingestion, storage, transformation, orchestration, governance, observability, semantic or serving layers, and downstream consumption.
04Why is data governance part of architecture?
Why is data governance part of architecture?
Ownership, permissions, metadata, lineage, quality, classification, and retention affect how safely and reliably data can be used across the organisation.
05How should a data platform support AI?
How should a data platform support AI?
AI workloads need dependable access to authorised, well-described, sufficiently fresh data. The platform should support retrieval, feature creation, analytical datasets, APIs, governance, lineage, and other serving patterns according to the application.
Related
Continue the technical discussion
Start a conversation
Design a data platform around the workloads it needs to support.
Programmatic can help assess the existing data estate, define target architecture, modernise pipelines, introduce governance and observability, and build data products for analytics and AI.