By Yash Amin, Founder and CEO · Published · Updated

    Back to Case Studies

    How we replaced one-off pipelines at an e-commerce company with a self-service data platform

    The problem wasn't a lack of data. It was the lack of a repeatable way to move, test and deploy it, so every request landed on one data team.

    Industry
    Retail and e-commerce
    Platform
    Modern analytical data platform
    Engagement
    Data platform modernization
    Our role
    Data platform engineering
    Self-service

    Teams build on a standard path. Every request used to go through the data team.

    CI/CD

    Automated testing and deployment, which were manual before

    Earlier

    Data problems caught by contracts and validation, not after reports broke

    Visible

    Pipeline health, freshness and lineage tracked, not investigated by hand

    About the Client

    A growing e-commerce organization where data drives customer experience, operations, reporting and its analytics and AI plans.

    Its systems include a website and mobile app, ERP, CRM, payments, inventory and warehouse platforms, marketing tools, marketplace channels and customer support.

    DetailSpecifics
    Industry and sizeRetail and e-commerce, growing organization
    Data environmentWebsite and mobile app, ERP, CRM, payments, inventory and warehouse, marketing, marketplaces and customer support
    PlatformModern analytical data platform
    TeamA central data engineering team that built and maintained the pipelines for every request
    State before engagementFragmented integrations, custom pipelines and manual workflows

    The Situation

    As the organization grew, every new report, integration or data request meant the central data team had to build another custom pipeline.

    Those workflows had grown out of scripts, scheduled jobs, APIs and manual steps. They were hard to standardize, hard to maintain and harder to scale, and the request backlog kept growing.

    The business also wanted stronger analytics and AI. That needed a data foundation that was reliable, observable and repeatable first.

    Challenges

    Fragmented data

    Business data sat across many operational systems with no consistent analytical foundation.

    One-off pipelines

    Each new requirement meant another custom integration, and the maintenance load grew with every one.

    A growing backlog on one team

    Engineering and business teams waited on the central data team for every new workflow.

    Problems found late

    Data quality issues and upstream schema changes surfaced after pipelines had run and reports were already affected.

    Manual testing and deployment

    Releases needed hands-on steps, and a failed pipeline took manual investigation to explain.

    A data foundation not ready for AI

    Analytics and AI plans depended on data that was reliable and accessible, and it was neither yet.

    Diagnosis

    The constraint was not the number of data sources. It was that no standard engineering path existed for building and running data workflows. Four causes sat behind it.

    No standard way to build a pipeline

    Scripts, scheduled jobs, APIs and manual steps had piled up over time. Each pipeline was built differently, so none could be reused.

    No contract between producers and consumers

    Upstream schema changes broke downstream workflows because nothing defined what a source had to deliver.

    Validation and testing came after the fact

    Limited data quality checks meant bad data moved through the pipeline before anyone saw it.

    No visibility into workflows

    When data changed unexpectedly, finding the cause meant reading scripts and asking the people who wrote them.

    The organization didn't need more pipelines. It needed a repeatable platform and a path for building them the same way every time.

    What We Did

    1. Standardized data integration

      Defined reusable patterns for APIs, databases, CDC, event streams, file ingestion and SaaS connectors. New sources now plug into an existing pattern instead of starting from scratch.

    2. Built the pipeline foundation

      Set up repeatable ETL and ELT patterns for ingestion, transformation and analytical workloads, landing in a central analytical warehouse that feeds curated business data products.

    3. Introduced the Golden Path

      Created one engineering workflow covering pipeline templates, data contracts, validation, automated testing, CI/CD and observability. A developer or business team follows it instead of raising a request.

    4. Built reliability into the workflow

      Data quality checks, orchestration, monitoring and documentation became steps in the path, not separate activities done after launch.

    Architecture: Before and After

    DharmOps e-commerce data platform architecture: fragmented systems feeding a central data team, the target platform with ingestion, processing, warehouse, data products and consumption layers, the Golden Path from manual to self-service, the order data flow, and before and after business impact

    Swipe the diagram sideways to read it.

    View diagram full size

    Before

    Business systems feed a central data team that builds and maintains every pipeline by hand. Custom integrations, inconsistent data and limited visibility pile up into a request backlog.

    After

    Standard ingestion, processing, quality controls and contracts feed one warehouse. Curated data products serve dashboards, self-service analytics and AI use cases through the Golden Path.

    Implementation Timeline

    1. Platform assessment

      Mapped systems, pipelines, integrations, deployment processes, data quality issues, bottlenecks and ownership.

    2. Golden Path design

      Defined reusable patterns for ingestion, transformation, validation, testing, deployment, monitoring and documentation.

    3. Pilot build

      Applied the pattern to one representative data workflow: integration, transformation, validation, testing, deployment and monitoring.

    4. Handover and documentation

      Documented the platform and trained the engineering team to operate it.

    5. Expansion

      Extended the proven patterns to more data products and domains.

    Before and After

    StepBeforeAfter
    Data requestsRaised with the central data teamFollowed through the standard self-service path
    Pipeline developmentCustom, one-off pipelinesReusable Golden Path templates
    TestingManual and inconsistentAutomated validation and testing
    DeploymentManualCI/CD
    Data qualityIssues found lateContracts and validation catch them early
    VisibilityLimitedMonitoring, lineage and alerts

    Results

    MetricBeforeAfterChange
    Data deliveryWaits on the central teamStandard engineering pathSelf-service
    Pipeline developmentOne-offReusable patternsRepeatable
    DeploymentManualCI/CDAutomated
    Data qualityDetected lateValidation built into workflowsCaught earlier
    VisibilityLimitedObservability and alertsTracked
    AI readinessBlocked by data cleanupGoverned analytical foundationReady to build on

    Tradeoffs and What We'd Do Differently

    1. Start with the actual bottleneck

      Begin with the constraint the organization is living with, not with a choice of tools.

    2. Prove one Golden Path first

      A single pilot workflow shows what works before the pattern is applied across every domain. It costs time up front and slows the first rollout.

    3. Treat Golden Paths as products

      A path only stays useful with an owner, documentation, support and regular improvement.

    4. Fix the data foundation before scaling AI

      Analytics and AI projects move faster once the underlying data is reliable, accessible and monitored.

    Technology Stack

    LayerTool
    Data engineeringPython, SQL, dbt
    IntegrationAPIs, CDC, Kafka and event streaming, database and SaaS connectors
    OrchestrationAirflow, Dagster
    Data platformSnowflake
    Platform engineeringDocker, Kubernetes, Terraform
    CI/CDAutomated testing, build and deployment
    OperationsLogs, metrics, alerts, pipeline monitoring
    Data qualitySchema, freshness, integrity and anomaly checks

    Data Platform Engineering Questions

    What is a Golden Path in data platform engineering?

    A Golden Path is one standard engineering workflow for building and running data pipelines. In this engagement it covered pipeline templates, data contracts, validation, automated testing, CI/CD and observability, so a developer or business team follows the path instead of raising a request with the central data team.

    Why do custom data pipelines become a bottleneck?

    Each new report or integration needed another custom pipeline built by one central team. Those pipelines grew out of scripts, scheduled jobs, APIs and manual steps, so they were hard to standardize, hard to maintain and impossible to reuse, and the request backlog kept growing.

    What are data contracts and why do they matter?

    A data contract defines what a source system must deliver. Without one, upstream schema changes broke downstream workflows because nothing stated what a source had to provide. Contracts plus validation catch those problems before pipelines run, not after reports are affected.

    How do you roll out a Golden Path without disrupting teams?

    Prove it on one representative workflow first: integration, transformation, validation, testing, deployment and monitoring. Once the pilot shows what works, extend the patterns to more data products and domains. The tradeoff is time up front and a slower first rollout.

    Why fix the data foundation before scaling AI?

    Analytics and AI plans depend on data that is reliable, accessible and monitored. Here, those plans were blocked by data cleanup until a governed analytical foundation with contracts, validation and observability was in place.

    How do you implement data contracts?

    Define what each source must deliver, then enforce it in the pipeline. Here, contracts, validation and automated testing became steps in the Golden Path, run through CI/CD. A source that changes its schema now fails a check before the pipeline runs, instead of breaking downstream reports afterward.

    What does a self-serve data platform architecture include?

    In this e-commerce platform: standard ingestion patterns for APIs, databases, CDC, event streams, file ingestion and SaaS connectors; ETL and ELT patterns landing in a central Snowflake warehouse; curated data products on top; Airflow or Dagster orchestration; and the Golden Path of templates, contracts, validation, CI/CD and observability that teams follow instead of raising requests.

    Is your data team the bottleneck?

    On a diagnostic call we walk through your current data landscape, pipeline bottlenecks and integration approach, and where a self-service path would remove the most requests.

    Start with a Diagnostic Call

    Tell us what's breaking.

    One call to walk through the symptoms. You'll leave knowing where to look first.

    Talk to Us

    Or email contact@dharmops.com