How to Build a Data Science Function From Scratch
A practical guide to standing up a data science function from zero — who your first hire should report to, whether to start with a data scientist or analytics engineer, and how to sequence early wins before scaling the team.

Most scale-ups don't fail at data science because the models are wrong. They fail because the function was never designed — a data scientist gets hired into a vacuum, reports to whoever won the budget argument, and spends the first six months arguing about where the data even lives. This guide covers the practical decisions: reporting lines, first hires, tooling, and how to sequence wins before you scale up the team.
What does "building a data science function from scratch" actually mean?
Building a data science function from scratch means establishing the people, reporting structure, tooling, and workflows needed to turn raw data into decisions and products — before any of that infrastructure exists. It is a data science function is the organisational unit responsible for analytics, statistical modelling, and machine learning, distinct from software engineering and business intelligence, though it usually grows out of one or both. For most Australian scale-ups, this isn't a green-field exercise — there's already a data warehouse of some sort, a BI tool someone half-configured, and a backlog of requests nobody owns.
Who should your first data hire report to?
Your first data hire should report to a technical leader — typically the CTO, VP Engineering, or a fractional CTO — not to a commercial function like marketing or finance, even if their early work serves those teams. Reporting into engineering protects data infrastructure decisions from being made ad hoc by whichever department shouts loudest, and it keeps data quality and pipeline reliability treated as engineering problems, which they are.

This matters because the earliest failure mode in a new data function is capture: the first data scientist becomes an analyst embedded in one team, answering ad hoc questions, with no time or mandate to build the pipelines and models that create compounding value. A clear reporting line to a technical leader — supported by CTO advisory if you don't have one in-house yet — keeps the mandate broad enough to build a function rather than a queue.
Should you hire a data scientist or an analytics engineer first?
Most scale-ups should hire an analytics engineer or data engineer before a data scientist, because a data scientist without clean, accessible data will spend most of their time on data wrangling rather than modelling. An analytics engineer is a hybrid role that builds and maintains the transformation layer between raw data sources and analysis-ready tables, typically using SQL-based tools within the modern data stack.
The exception is when you already have decent data infrastructure — a working warehouse, reasonably reliable pipelines, existing BI dashboards — and the gap is genuinely analytical or predictive capability. In that case, a data scientist can be productive from day one. If you're not sure which situation you're in, an honest audit of your data infrastructure will tell you faster than guessing.
How should the team be structured as it grows?
There is no single correct structure, but most functions evolve through a recognisable sequence: centralised first, then either embedded or hub-and-spoke as demand diversifies across business units. The right model at 3 people is rarely the right model at 15.
| Model | How it works | Best for | Main risk |
|---|---|---|---|
| Centralised | One team serves the whole company, prioritised centrally | Early stage, 1-5 people, limited demand | Becomes a bottleneck as requests multiply |
| Embedded | Data scientists sit inside product or business teams | Mature orgs with strong domain-specific needs | Duplicated tooling, inconsistent standards |
| Hub-and-spoke | Central team owns platform/standards; embedded members execute locally | Growing orgs, 15+ people, multiple business units | Requires strong central leadership to avoid drift |
For most companies in the 50-2000 employee range, starting centralised and moving to hub-and-spoke as the team passes ten to fifteen people is the least risky path. It avoids the premature complexity of managing distributed reporting lines before there's enough work to justify them.
What tooling should you choose early on?
Early tooling decisions should optimise for time-to-first-insight and low operational overhead, not for scale you don't yet need. A managed cloud data warehouse, a SQL-based transformation tool, and a BI layer are usually sufficient for the first 12-18 months — resist the urge to stand up a full MLOps platform before you have a model worth operationalising.

Buying a platform is not the same as building a function. Platform vendors can accelerate specific workflows, but as several data-tooling vendors themselves acknowledge, a platform still requires internal teams to operate it and doesn't solve the underlying talent and expertise gap on its own. The tooling should follow the team's maturity, not lead it. If your organisation has data assets but lacks internal data engineering or ML expertise to operate whatever you buy, that gap needs to be closed with people or a partner before the tooling investment pays off.
How do you sequence early wins before scaling the function?
Sequence early wins by picking one or two business-critical questions with existing, reasonably clean data, delivering an answer within weeks, and using that credibility to justify the next hire — rather than starting with an ambitious predictive model that takes months to prove out. Early wins should be visible to the executive team, not just useful to the requesting department.
A practical sequence looks like this:
- Audit what exists. Map current data sources, pipelines, and reporting before hiring anyone. You cannot sequence wins against a backlog you haven't measured.
- Fix the plumbing first. Even one broken or untrusted data source undermines confidence in everything built on top of it.
- Deliver a descriptive win. A reliable dashboard or report that replaces a manual, error-prone process is often more valuable early than a predictive model — it's visible, low-risk, and builds trust.
- Layer in one predictive or ML use case. Choose a problem with a clear business owner, measurable outcome, and tolerance for iteration. This is where AI product strategy work earns its keep — picking the right first use case matters more than picking the most technically interesting one.
- Only then scale the team. Use the credibility from steps 1-4 to justify the next two or three hires, and start formalising the reporting structure described above.
What are the most common mistakes when building this function?
The most common mistakes are hiring a senior data scientist before the data infrastructure exists, reporting the function into a commercial team that treats it as an internal agency, and investing in an MLOps platform before there's a production model to operate. All three come from the same root cause: sequencing the investment before the demand and infrastructure are ready to support it.
A related mistake is treating data science and application modernisation as unrelated workstreams. If your core systems are a legacy monolith with no clean data access layer, no analytics team — however well-resourced — can move quickly. In that situation, application modernisation work often needs to happen in parallel with, or even before, the data science hiring plan.
Building this function well takes longer than most roadmaps assume
Standing up a data science function from scratch is rarely a straight line. It requires sequencing infrastructure before headcount, headcount before tooling, and quick wins before ambitious models — and getting any of that order wrong tends to show up as frustrated hires and stalled projects six months later. For further reading on the infrastructure layer this all depends on, see our piece on data infrastructure: building the foundation for AI, and for a wider view of how AI fits into the roadmap, browse our insights.
If you're planning your first data science hires and want a second opinion on sequencing, reporting lines, or tooling before you commit budget, get in touch — we're happy to talk through what's realistic for your stage, even if the answer is "hire an analytics engineer before a data scientist."
Chris Kerr
Partner at Horizon Labs, an AI product consultancy and venture studio. A commercially focused product and technology leader with 20+ years building and scaling digital platforms, teams, and businesses across SaaS, travel, eCommerce, logistics and transport, and digital marketing — operating at the intersection of product, engineering, and data. Writes about platform strategy, AI transformation, modern data ecosystems, and the operational discipline that separates AI demos from AI products.


