How long does it take to build a production-ready AI data platform?

    The honest timeline for a mid-market Australian business is 12 to 20 weeks, not six. Here's where the weeks actually go, and which of them you can genuinely compress.

    By Will Turner · Founder & Head of Data & AI25 April 20267 min read

    For a mid-market Australian business with a managed cloud warehouse, a tolerable amount of existing data chaos, and a single well-defined first AI or analytics use case, the honest timeline to a production-ready data platform is 12 to 20 weeks. Not six. Not six months. Twelve to twenty.

    Critical caveat, because this is the part most vendors hide: that range assumes you're buying the core platform (Snowflake, BigQuery, Databricks) rather than building your own on open-source components. If you insist on building the infrastructure layer yourself (Kubernetes, your own orchestration, your own table-format catalog, your own RBAC, your own observability), you are adding six to nine months to the first production use case, sometimes more. The 12-to-20-week number is the managed-platform number. The build-your-own number is another article, and the short answer is don't. Every business that's shipped faster than 12 weeks on a managed platform either cut corners that came back to hurt them, or didn't actually ship production-ready (they shipped a demo).

    Here's where the weeks go.

    Phase 1: discovery and design (weeks 1–3)

    • Data source inventory. Every system the platform needs to ingest from, what state it's in, who owns it, how change-data-capture works (or doesn't).
    • Use case shaping. The specific decision, the specific metric, the specific dashboard or model the first release unlocks. This is where 40% of vague "AI platform" projects die, because nobody can answer what the thing is actually for.
    • Platform selection. If not already decided. Usually a week, occasionally six if there are procurement hurdles.
    • Access and security design. Who sees what. SSO integration. Row-level and column-level access controls. Data residency if regulated.
    • Naming and modelling conventions. Not exciting, not optional. Decisions here lock in for five years.

    Cheap to rush, expensive to get wrong. Budget 3 weeks, accept 2 if the scope is narrow and the owner is decisive.

    Phase 2: foundation (weeks 4–9)

    • Platform provisioning. The warehouse, the ingestion tool, dbt, orchestration, observability, cost monitoring, Git integration, CI/CD. Each individually is a day; collectively the integration work takes two to three weeks.
    • First ingestion pipelines. The three to five most important source systems. Plus a skeleton for the rest that'll come later.
    • Staging layer. Raw tables cleaned, typed, deduped. Nothing opinionated yet. Just "we can trust what's in here."
    • Monitoring and alerting. Freshness SLAs, quality alerts, cost caps. This is where cheap projects cut corners and where expensive incidents start in month three.

    This phase takes six weeks of good focused work. If you're hearing "three weeks" in a vendor's plan, either they're assuming an existing foundation or they're skipping monitoring.

    Phase 3: first use case end-to-end (weeks 10–15)

    • Transformations. Source-of-truth models for the domain the first use case needs. Tested, documented, reviewed.
    • Semantic definitions. "What is revenue." "What is a customer." Written once, stored in code, reviewed by the business.
    • Serving layer. Dashboard, BI tool integration, API endpoint, or ML feature store (depending on what the use case needs).
    • User acceptance testing. With the actual humans who'll act on the outputs. This is the step that teams hope is quick and is never quick.
    • Training and enablement. The people who'll use the platform on Monday need to be able to use it without asking the data team every question.

    Six weeks is a realistic middle. Two weeks is possible for a simple use case on clean data (rare). Ten weeks is common when the first use case turns out to be three use cases wearing a trench coat.

    Phase 4: harden and hand over (weeks 16–20)

    • Performance tuning. The model or dashboard that was fine on 90 days of data needs to work on 3 years.
    • Cost optimisation. Query patterns the team wrote in week 4 are almost certainly expensive. Fix before bad habits become infrastructure.
    • Runbooks. On-call procedures, escalation paths, known-failure modes. Who restarts the pipeline when it breaks at 3am.
    • Documentation. The boring, essential kind. What every table is, who owns it, how it's loaded, what to do when it looks wrong.
    • Handover. From the build team to the ops team, whether those are different people or different hats on the same people.

    This phase exists in every serious platform and doesn't exist in any platform that later fails in month six. The correlation is not an accident.

    What makes it faster

    • Using a managed platform, not an open-source stack. This is by a wide margin the largest lever. Building your own platform layer adds 6 to 9 months to first production use case, sometimes a year. There is no situation in the mid-market where the DIY path produces a production-ready platform in less time than Snowflake, BigQuery, or Databricks does. Businesses that think there is are still in month eight of month-four estimates.
    • Only one first use case in scope. Easy 2-4 week saving. Businesses that insist on three first use cases burn the saving and then some.
    • Clean, API-first source systems. Rare in the mid-market, worth a lot when true. Dirty or API-hostile source systems add 3-5 weeks that look like data quality work but are really integration work.
    • An engaged executive sponsor who answers questions the same day and makes decisions in meetings. Invisible to the schedule until it's absent, at which point it's worth 3-5 weeks of delay across the project.

    What makes it slower

    • Data quality is worse than advertised. The #1 cause of timeline overruns. Nobody knows how bad the data is until the build team tries to join it, and then it's always worse than the walkthrough implied.
    • Source system access takes longer than expected. Security reviews, network rules, service account provisioning. Budget 2–4 weeks for this, plan like it's one week and you'll eat the difference.
    • The first use case grew during build. Scope creep in analytics and AI work looks like "while we're here, let's add…" Every addition is a week you didn't schedule.
    • Decisions get batched at the sponsor's weekly meeting. Each blocker is a week.

    The honest version of "we need this fast"

    If a sponsor insists on 8 weeks for a production-ready platform, there are only three honest replies:

    1. We can do it in 8 weeks if the first use case is narrow, the data is clean, and the platform is managed. If any of those aren't true, 8 weeks is 14 weeks pretending.
    2. We can do it in 8 weeks by cutting monitoring, testing, and documentation. The platform will be in production and it will break in month three, and the break will feel worse than the delay you were trying to avoid.
    3. We can get a demo running in 8 weeks that looks like the real thing. It'll be fine for the board meeting. We'll still need 10 more weeks to make it production-ready.

    The teams that pick option 1 and scope honestly tend to ship. The teams that pick option 2 show up in our stalled-project queue. The teams that pick option 3 at least go in with their eyes open.

    The single best cost-saver

    Start the platform with one use case, not five. One use case pulls the platform into existence in twelve weeks. Five use cases keep the platform in a design phase for six months.

    Common Questions

    Frequently asked

    Got one of these problems in front of you?

    Beyond Data runs engagements that put the ideas in this insight into practice.

    Cookie preferences

    We use essential cookies to run the site. With your permission, we also use analytics cookies to understand what content helps visitors make better data and AI decisions.