Skip to content

Data Architect Agent: Agent Workflow & Artifacts

The Data Architect Agent is the strategic design engine within the Datapunkt ecosystem, purpose-built to plan, model, and deliver enterprise-grade data architectures in minutes instead of months. It eliminates the slow, high-cost cycles of traditional consulting by generating production-ready High-Level Designs (HLDs), Low-Level Designs (LLDs), and deployable Terraform Foundation scripts across seven industry-proven architectural paradigms.

This guide walks you through the complete step-by-step operational workflow of using the Data Architect Agent, defines every core output artifact, and explains the benefits of integrating it into your daily architectural practice.

first

Supported Architecture Patterns

The Data Architect Agent natively supports seven enterprise-grade data architecture paradigms. Each pattern is available as a first-class design target, and the agent can combine multiple patterns within a single blueprint when your business requirements demand it.

MEDALLION

The Medallion Architecture organizes data into progressive refinement layers, typically Bronze (raw ingestion), Silver (cleansed and conformed), and Gold (business-ready aggregates). The Data Architect Agent generates complete layer definitions with explicit schema contracts, data quality checkpoints, and inter-layer transformation rules for every stage of the pipeline:

  • Bronze Layer: Raw, unmodified ingestion tables that preserve the exact structure of source systems. The agent defines landing zone schemas, file format specifications (Parquet, Delta, Avro), partitioning strategies based on ingestion timestamps, and retention policies that comply with regulatory requirements.
  • Silver Layer: Cleansed, deduplicated, and type-cast tables with enforced naming standards. The agent generates deduplication logic, type coercion rules, null-handling strategies, and slowly changing dimension (SCD) merge patterns to maintain historical accuracy.
  • Gold Layer: Business-ready analytical tables, denormalized aggregates, and metric-specific views. The agent designs star schemas, wide tables, or domain-specific data marts optimized for query performance and direct consumption by BI platforms and machine learning pipelines.

SEMANTIC

The Semantic Layer Architecture introduces a business-meaning abstraction between raw physical tables and end-user consumption. The Data Architect Agent generates comprehensive semantic model definitions that translate complex joins, calculations, and business rules into reusable, governed metric definitions:

  • Metric Definitions: The agent creates formalized metric catalogs with explicit calculation logic, grain specifications, dimension hierarchies, and time-intelligence patterns that guarantee consistent results across all downstream consumers.
  • Business Glossary Mappings: Every physical column in the target schema is mapped to a business-friendly name, description, data steward, and lineage reference, ensuring that analysts and stakeholders interact with meaningful labels rather than cryptic database identifiers.
  • Access Control Layers: The agent defines role-based access policies at the semantic layer, restricting visibility to specific metrics, dimensions, or row-level segments based on user groups, departments, or regulatory classifications.

DATA VAULT

Data Vault 2.0 is an enterprise modeling methodology designed for auditability, historical tracking, and agile iteration. The Data Architect Agent generates complete Data Vault structures including Hubs, Links, Satellites, and Point-in-Time (PIT) tables:

  • Hubs: Business key entities with hash keys, load timestamps, and record source identifiers. The agent calculates deterministic hash keys using MD5 or SHA-256 algorithms and ensures global uniqueness across federated source systems.
  • Links: Many-to-many relationship tables connecting two or more Hubs. The agent generates composite hash keys from participating Hub business keys and includes degenerate keys where applicable.
  • Satellites: Temporal attribute storage tables that capture every historical change to a Hub or Link. The agent generates hash-diff columns for efficient change detection, defines load-end-date patterns for bitemporal tracking, and structures effective-dating logic for point-in-time reconstruction.
  • PIT Tables: Performance acceleration structures that pre-join the latest Satellite records to their parent Hubs, enabling fast analytical queries without expensive temporal lookups.
  • Bridge Tables: Pre-computed link traversal structures that flatten complex relationship paths for reporting consumption.

DATA MESH

Data Mesh is a decentralized organizational architecture where autonomous domain teams own, produce, and serve their data products. The Data Architect Agent generates domain-oriented blueprints with self-serve infrastructure templates, federated governance contracts, and inter-domain data product interfaces:

  • Domain Boundaries: The agent maps your organizational structure to logical data domains, each with an independent data product catalog, schema registry, and quality SLA contract.
  • Data Product Specifications: For each domain, the agent generates a formal data product definition including input ports (source contracts), output ports (published interfaces), transformation logic, SLA guarantees (freshness, completeness, accuracy), and discoverability metadata.
  • Federated Governance: The agent produces a global governance overlay that enforces naming conventions, security classifications, PII handling rules, and interoperability standards across all autonomous domains without centralizing ownership.
  • Self-Serve Platform Templates: Modular Terraform templates that domain teams can instantiate independently to provision their own storage, compute, and orchestration infrastructure within guardrails defined by the central platform team.

KAPPA

The Kappa Architecture eliminates batch processing entirely, using a single real-time streaming pipeline as the canonical data processing path. The Data Architect Agent generates streaming-first blueprints with event log retention, stream processing topologies, and materialized view definitions:

  • Event Log Design: The agent defines immutable, append-only event logs with configurable retention periods, compaction strategies, and partition key layouts optimized for high-throughput streaming engines like Apache Kafka, Google Pub/Sub, or Amazon Kinesis.
  • Stream Processing Topologies: Complete processing graph definitions including windowing strategies (tumbling, sliding, session windows), watermark configurations for late-arriving data, and exactly-once processing guarantees.
  • Materialized Views: The agent generates continuously updated materialized views that serve analytical queries directly from the streaming layer, eliminating the need for batch ETL entirely.
  • Reprocessing Patterns: When business logic changes, the agent designs replay topologies that reprocess the entire event log through the updated pipeline version, producing corrected materialized views without data loss.

LAMBDA

The Lambda Architecture maintains parallel batch and real-time processing paths that merge into a unified serving layer. The Data Architect Agent generates dual-path blueprints with batch layer definitions, speed layer configurations, and merge logic for the serving layer:

  • Batch Layer: Scheduled, high-throughput processing pipelines that operate on complete historical datasets. The agent defines batch job schedules, input/output format specifications, incremental processing checkpoints, and fault tolerance configurations.
  • Speed Layer: Low-latency streaming pipelines that process individual events in real time. The agent generates stream processing definitions with micro-batch intervals, state management configurations, and exactly-once delivery guarantees.
  • Serving Layer: The merge point where batch and speed outputs are combined into a single, queryable view. The agent designs the merge logic, conflict resolution rules (batch-wins vs. speed-wins), cache invalidation strategies, and query routing configurations.
  • Consistency Reconciliation: The agent generates reconciliation jobs that periodically compare batch and speed layer outputs to detect and correct drift, ensuring long-term consistency between the two processing paths.

LAKEHOUSE

The Lakehouse Architecture unifies data lake storage economics with data warehouse query performance, typically implemented using Delta Lake, Apache Iceberg, or Apache Hudi table formats. The Data Architect Agent generates Lakehouse blueprints with ACID-compliant table definitions, schema evolution strategies, and unified batch-streaming processing configurations:

  • Open Table Formats: The agent generates table definitions using Delta Lake, Iceberg, or Hudi with ACID transaction guarantees, time-travel capabilities, and schema evolution support. Each table includes explicit partition pruning specifications, file compaction schedules, and Z-ordering configurations for optimal query performance.
  • Unified Batch-Streaming Ingestion: The agent designs ingestion pipelines that support both batch file drops and real-time streaming events into the same Lakehouse table, using merge-on-read or copy-on-write strategies depending on the workload profile.
  • Query Engine Compatibility: Generated schemas are optimized for direct consumption by Spark SQL, Trino, Presto, Databricks SQL, and Snowflake external tables, ensuring maximum flexibility across your analytics toolchain.
  • Storage Optimization: The agent defines file compaction jobs, vacuum schedules, and statistics collection routines to maintain optimal read performance as table sizes grow into petabyte scale.

Operational Workflow

The Data Architect Agent utilizes a dual-pane workspace designed to streamline complex architectural planning while maintaining strict security boundaries around your infrastructure metadata.

Step 1: Subscribing to the Agent

To start using the agent, you must obtain a subscription through the platform marketplace:

  1. Log in to your Agentpunkt platform account.
  2. Navigate to the platform marketplace catalog and search for the Data Architect Agent.
  3. Select the subscription plan that fits your team's needs (e.g., Weekly, Bi-weekly, or Monthly).
  4. Once subscribed, the agent appears on your Hired Agents Page, ready to be activated.

Step 2: Workspace Session Initialization

Open a new workspace session to establish your sandboxed architectural environment:

  1. Click "Start Session" next to the Data Architect Agent on the Hired Agents Page.
  2. The user interface splits into two primary panes:
    • Left Pane (Interactive Chat & Logs): Here you communicate with the agent in natural language, describe your business objectives, paste existing schema definitions, specify cloud platform preferences, and monitor the agent's step-by-step reasoning and design decisions.
    • Right Pane (Workspace & Blueprint Preview): This pane displays generated High-Level Designs, Low-Level Designs, Terraform Foundation scripts, schema diagrams, and ready-to-deploy configuration files.

Step 3: Defining Your Business Context and Architecture Pattern

Establish the foundational parameters that will drive the entire design:

  1. In the Left Pane, describe your business objectives, expected data volumes, target cloud platform, existing systems, and compliance requirements.
    • Example prompt: "We are a fintech company building a real-time fraud detection platform on GCP. We expect 80 million events per day from payment gateways, mobile apps, and partner APIs. We need to comply with GDPR and PCI-DSS. We want a LAMBDA architecture with a speed layer for real-time scoring and a batch layer for model retraining. Our analytics team uses BigQuery and Looker."
  2. Specify which architecture pattern or combination of patterns you want the agent to use. The agent supports MEDALLION, SEMANTIC, DATA VAULT, DATA MESH, KAPPA, LAMBDA, and LAKEHOUSE as first-class design targets.
    • Example prompt: "Use a LAKEHOUSE architecture with Delta Lake table format for our core storage layer, overlay a MEDALLION progression from Bronze through Gold, and add a SEMANTIC layer for our business metrics catalog."
  3. The agent confirms your selections, asks clarifying questions about edge cases, and begins generating the architectural blueprint.

Step 4: High-Level Design (HLD) Generation

The agent produces a strategic overview of your new data landscape:

  1. The HLD document appears in the Right Pane, containing:
    • Conceptual data models showing entity relationships and data domain boundaries.
    • Platform selection rationale with explicit justifications for each cloud service, table format, and processing engine.
    • High-level integration points mapping data flows between business units, source systems, and analytical consumers.
    • Governance overlay showing compliance zones, PII boundaries, and access control hierarchies.
    • Architecture pattern application showing exactly how your selected patterns (e.g., MEDALLION + LAKEHOUSE) interact and complement each other.
  2. Review the HLD, request modifications to domain boundaries, platform choices, or governance rules, and iterate until the strategic design meets your requirements.

Step 5: Low-Level Design (LLD) Generation

Once the HLD is approved, request the detailed technical specifications:

  1. In the Left Pane, ask the agent to generate the LLD.
    • Example prompt: "Based on our approved HLD, please generate a comprehensive Low-Level Design including target schemas for all layers, partition strategies, clustering keys, and detailed pipeline execution sequences."
  2. The LLD document appears in the Right Pane, containing:
    • Complete DDL scripts for every table in every layer (Bronze, Silver, Gold, Hubs, Links, Satellites, etc.).
    • Partition key definitions and clustering column specifications for each table.
    • Schema mapping definitions showing exact column-level transformations between layers.
    • Pipeline execution sequences with dependency graphs, scheduling intervals, and resource sizing recommendations.
    • Index definitions, materialized view specifications, and query optimization hints.
    • Data quality checkpoint definitions embedded at each layer boundary.

Step 6: Terraform Foundation Generation

Deploy your designed architecture to your cloud platform automatically:

  1. Once you are satisfied with the logical design, request the infrastructure code.
    • Example prompt: "Generate the modular Terraform files to provision this LAKEHOUSE architecture on GCP, including Cloud Storage buckets for the Bronze layer, BigQuery datasets for Silver and Gold, Dataproc clusters for Spark processing, IAM service accounts with least-privilege access, and Pub/Sub topics for the streaming ingestion layer."
  2. A complete set of .tf modules appears in the Right Pane, organized by infrastructure component:
    • Networking (VPC, subnets, firewall rules)
    • Storage (buckets, datasets, tables)
    • Compute (Dataproc, Dataflow, Cloud Functions)
    • Identity (IAM roles, service accounts, least-privilege policies)
    • Monitoring (alerting rules, log sinks, SLA dashboards)
  3. Review the generated scripts, configure your environment variables, and run terraform apply to provision your entire architecture stack.

Step 7: Artifact Review and Export

Once the agent completes any workflow, review and export your deliverables:

  1. All generated files are displayed in the Right Pane with syntax highlighting and inline documentation.
  2. Download individual files or export the entire session as a compressed archive.
  3. Copy specific code blocks directly from the Right Pane into your local repository, CI/CD pipeline, or version control system.

Benefits: What Makes It Good?

  • Architectural Acceleration: Shrinks enterprise data architecture design from weeks or months of workshops, whiteboard sessions, and manual documentation to under one hour of interactive agent-guided design, recovering up to 95% of senior architect capacity.
  • Seven First-Class Patterns: Native support for MEDALLION, SEMANTIC, DATA VAULT, DATA MESH, KAPPA, LAMBDA, and LAKEHOUSE means the agent adapts to your business requirements rather than forcing you into a single rigid methodology.
  • End-to-End Deliverables: Every session produces a complete artifact chain from strategic HLD through detailed LLD to deployable Terraform Foundation scripts, eliminating the gap between design and implementation.
  • Pattern Composition: The agent intelligently combines multiple architecture patterns (e.g., MEDALLION + LAKEHOUSE + SEMANTIC) within a single coherent blueprint, handling the complex interactions between patterns that typically require deep specialist knowledge.
  • Production-Ready Infrastructure: Generated Terraform modules include least-privilege IAM policies, network isolation, retention policies, monitoring hooks, and governance labels, making them safe to deploy directly to production cloud environments.
  • Governance by Design: Every generated blueprint includes built-in compliance zones, PII boundary definitions, access control hierarchies, and regulatory retention policies, ensuring your architecture is secure and compliant from day one.

Contact Us