Engineering Management
From Zero to Scale: Seven Principles for Building a High-Performing Data & AI Organisation
What I learned building a Data Engineering and ML practice from zero to more than sixty professionals across a complex, regulated, and multicultural environment.
Executive brief
A high-performing Data & AI organisation is built by translating business ambition into repeatable capabilities, clear decision rights, strong leadership, explicit accountability, usable governance, inclusive operating practices, and a balanced measurement system. Scale is real when the organisation produces trusted, governed, and measurable outcomes without depending on individual heroes.
Key takeaways
- Define repeatable business capabilities before drawing the organisation chart.
- Judgement, ownership, communication, learning, architecture trade-offs, and collaboration outlast tool-specific fluency.
- A leadership layer becomes real when decisions continue without accumulating around one person.
- Autonomy works when outcomes, risks, ownership, and escalation paths are explicit.
- Good governance reduces decisions, accelerates delivery, and generates evidence as a by-product of work.
- Multicultural teams need written decisions, accessible language, procedural escalation, and end-to-end ownership.
- Organisational value must be read across a balanced set of delivery, people, client, quality, and maturity signals.
Building a Data & AI organisation is not primarily a hiring exercise. It is an exercise in translating business ambition into capabilities, decision rights, standards, leadership behaviours, and repeatable delivery.
TL;DR
- Start with the outcomes the business must produce repeatedly, not a proposed organisation chart.
- Hire for judgement, ownership, communication, learning, architecture trade-offs, and collaboration—not only technology keywords.
- Build the leadership layer before decision latency makes the need obvious, and delegate decisions rather than tasks alone.
- Combine freedom in execution with clear outcomes, visible risks, and one named owner for every deliverable.
- Treat governance as reusable delivery infrastructure that reduces ambiguity and produces evidence continuously.
- Design explicitly for different languages, time zones, and working cultures; written decisions are part of the operating model.
- Read growth, utilisation, retention, predictability, quality, succession, client outcomes, and capability maturity together.
When I began building a Swiss and DACH Data Engineering and ML capability, there was no mature organisation waiting to be scaled. We had to define the roles, build credibility with clients, establish technical and delivery standards, recruit carefully, develop leaders, and create a commercial model that could sustain growth. Over several years, that capability grew from zero to more than sixty professionals. The most important lessons were not about technology. They were about people, trust, accountability, and organisational design.
What follows is not a maturity model, and it is not a sequence. The seven principles below were learned in a fairly disorderly way, several of them by getting the first attempt wrong, and they interact more than any list can properly show. I have tried to write them as they were actually encountered rather than as they would appear in a framework.
1. Start with the business mandate, not the org chart
The first artefact most people reach for when asked to build a capability is an organisation chart. It is a comfortable artefact. It has boxes, it has a headcount total, and it can be presented to a steering committee inside ten minutes. It is also, in my experience, close to useless as a starting point, because it encodes an answer to a question that has not yet been asked.
The question worth asking first is what the organisation needs to be able to do repeatedly, at what quality, under whose supervision, and on what timescale. In a regulated environment that question has an unglamorous shape: the business needs to be able to bring a new data source into a governed platform without a six-week negotiation about ownership; it needs to be able to demonstrate lineage from a reported figure back to its source without reconstructing it by hand; it needs to be able to put a model into production and keep it there, with someone accountable for its behaviour after the launch email has been sent. Those are capability statements. Roles follow from them. Headcount follows from roles and from the shape of demand. The chart is the last thing you draw, and it is best understood as a record of decisions already taken about outcomes and ownership rather than as a plan.
The industry has an instructive memory on this point. Through the middle of the last decade a great many organisations hired modelling capacity well ahead of engineering capacity, on the reasonable-sounding theory that the scarce, valuable skill was the modelling. A significant proportion of those organisations then spent the following two years discovering that most of the actual work was access, ingestion, quality, and the slow business of establishing what a customer record meant across three systems. Some of the most capable people I have interviewed since then came out of those teams, and they arrived with a healthy scepticism about capability plans that begin at the interesting end.
Demand shape deserves more attention than it usually gets, particularly where the capability has to sustain itself commercially. Headcount without demand becomes a bench problem, which is a financial problem with a morale problem attached to it. Demand without headcount becomes a credibility problem, and credibility, once lost with a client, is recovered on a much longer timescale than a hiring cycle. The only reliable way I found through this was to hire ahead of contracted demand exclusively where the capability took longest to build—senior data engineering and architecture, essentially—and to hold everything else close to the demand curve. That judgement was not always right. At one point we built out a specialism that the local market was not yet willing to pay for at the rate we needed, and we absorbed that for the better part of a year before redeploying the people into work that was closer to what clients were actually buying.
2. Hire for judgement, not only technology keywords
Over the years I have conducted several hundred technical interviews. The most consistent lesson from them is that fluency and judgement are easily confused, and that fluency is by far the easier of the two to demonstrate in an hour.
Keyword hiring is seductive because it is fast and defensible. A requisition asks for a particular stack, the CV shows the stack, the candidate can discuss the stack, and the decision feels evidence-based. What it does not test is what happens when the requirement changes in the third week, or when the source system turns out to behave differently from its documentation, or when a business stakeholder asks a question that the architecture cannot currently answer. Tooling knowledge depreciates on a schedule that has, if anything, accelerated. Judgement compounds.
In practice I came to assess six things, and to assess them by observation rather than by asking about them directly, because people are well rehearsed at describing their own strengths.
Problem-solving shows up most clearly when the problem is deliberately under-specified. What matters is not the answer but what the candidate does in the first two minutes: whether they ask what the data is for, who consumes it, what happens when it is wrong, and what the tolerance for latency is—or whether they begin designing immediately.
Ownership is audible in how someone narrates something that went badly. Accounts of failure that contain no first person, and in which the causes are all upstream, are informative.
Communication is tested by asking a candidate to explain a technical constraint to a non-technical audience, because in a regulated client environment a data engineer will spend a substantial part of their working life talking to risk, audit, compliance, and business owners. The ability to make a constraint intelligible to those audiences is not a soft skill but a delivery skill.
Learning ability is best approached through something the person has learned recently, and specifically how they know they have learned it.
Architecture judgement reveals itself in trade-offs the candidate accepted knowingly, and in whether they can name a decision they would now reverse; people who have never reversed anything have usually not owned anything for long enough.
Collaboration is most visible in how someone handles disagreement with a more senior colleague, which in some working cultures is a question about courage and in others a question about channel.
The hires I got wrong were, with fairly boring consistency, the ones where I over-weighted fluency. Someone articulate, well-prepared, and technically current is genuinely pleasant to interview, and it is easy to read that as capability rather than as presentation. The counter-measure I eventually adopted was procedural rather than intuitive: more than one interviewer, an explicit written assessment against each of the six dimensions before any discussion took place, and a rule that a strong signal on fluency alone was not sufficient. It did not eliminate the error. It reduced it.
3. Build a leadership layer early
At roughly fifteen people, one leader can still hold the whole picture—every client, every architecture decision, every person's development conversation. At thirty, that is no longer true, although it is still possible to pretend for a while, largely by working longer hours. At sixty the pretence fails, and it fails visibly, usually in the form of decision latency: work sitting in a queue waiting for one person's attention, escalations arriving late because the path was informal, and architects who have quietly stopped proposing things because they know the proposal will be re-decided elsewhere.
Building the leadership layer therefore has to begin well before the point at which it becomes obviously necessary, which means building it while the organisation still appears to be functioning without it. That is an uncomfortable investment to make and an easy one to defer.
Three things mattered more than the rest. The first is recognising that moving a strong engineer into a technical lead role is a change of job rather than a promotion within the same one, and that a good number of excellent engineers are made temporarily worse by it if the transition is unsupported. The second is delegating decisions rather than tasks. Delegating tasks while retaining the decisions produces exactly the queue described above—one server, unbounded arrivals—and it also communicates something about trust that no amount of encouragement will offset. The third is writing down decision rights: which architecture decisions sit with the architects, which staffing decisions sit with the leads, what requires commercial sign-off, and what the escalation path and its expected timeframe are. Written decision rights feel bureaucratic when the team is small and become the load-bearing structure when it is not.
Succession is the part most often treated as an aspiration rather than a control. The version that worked was concrete: for each role whose absence would materially disrupt delivery, name the person who could hold it for three months, and if no such person exists, treat that as an entry on the risk register rather than as an item for next year's development plan. It is a slightly deflating exercise the first time you run it.
The honest test of whether a leadership layer is real is what happens when the leader is unavailable for two weeks. If decisions accumulate and escalations wait, the layer is nominal regardless of what the chart says. I delegated architecture sign-off more slowly than I should have. It cost decision speed, and it told a group of capable architects something about my confidence in them that I would not have said out loud.
4. Combine autonomy with explicit accountability
The leadership philosophy I ended up with is compact enough to state in a sentence: clear outcomes, visible risks, defined ownership, and freedom in execution. Each of those four carries weight, and the failure modes come from dropping one of them rather than from getting the balance subtly wrong.
Autonomy without defined outcomes does not produce creativity; it produces divergence, because people who have not been told what success looks like will each infer something reasonable and the inferences will not match. Accountability without autonomy produces something worse, which is a competent person doing work they are not permitted to shape, and that is one of the more reliable predictors of a resignation six months later.
The mechanics are unremarkable. Every deliverable has a named owner, singular; shared ownership between two people is, operationally, ownership by neither. Risks are raised early and treated as information rather than as an admission, and this is the part that requires active maintenance, because the moment raising a risk costs the person who raised it something—in reputation, in scrutiny, in a difficult meeting—the risk register begins to describe a project that does not exist. In a regulated environment that failure is not merely inconvenient. Escalation is defined as a normal act with a route and a time expectation attached, so that using it is a procedure rather than a judgement about a colleague.
The pattern I observed repeatedly is that teams do not misuse autonomy when the outcome is unambiguous. They misuse it when the outcome was never properly defined and they were left to guess, and then the guess is treated afterwards as a discipline problem.
5. Treat governance as an enabler
Governance has a poor reputation among engineering teams, and a good deal of that reputation is earned. Much of what circulates under the name is written for an auditor rather than for the people doing the work, arrives as a document nobody reads, and functions as a tax collected at the end of delivery rather than as anything that helps during it.
Governance that works does five things. It reduces ambiguity, so that a team does not have to re-decide from first principles how environments are separated or how access is granted. It prevents the same mistake from being made repeatedly across teams that have no reason to know about each other's incidents. It supplies reusable standards, which is the only mechanism I know of that makes a second project cheaper than the first. It accelerates delivery, because most of what slows delivery in a regulated setting is not engineering difficulty but unresolved questions about permission and evidence. And it makes accountability fair, which is the benefit that gets discussed least and matters most: in the absence of written standards, quality judgements collapse into the personal taste of whoever happens to be reviewing, and performance conversations become impossible to conduct honestly.
The practical test I came to use is whether a piece of governance reduces the number of decisions a team has to make from scratch. If it adds decisions, it is overhead wearing governance as a costume.
Our first attempt failed this test comprehensively. It was a long standards document, thorough, defensible, and almost entirely unread. The version that worked was much shorter: a small set of non-negotiables—lineage captured as a by-product of the pipeline rather than reconstructed afterwards, environment separation, access control, testing applied to data as well as to code—each with a named owner and a worked example, with everything outside that set left to the teams. It was short enough to remember, which turned out to be the binding constraint. The lineage discipline in particular repaid itself in a way that is difficult to argue for in advance and obvious in retrospect: teams that produce evidence continuously spend their audit preparation reading rather than rebuilding.
6. Build a multicultural operating model
A capability that operates across several countries is not a single organisation with travel. It is a set of working cultures that have to interoperate across time zones, languages, professional traditions, and quite different assumptions about how disagreement is expressed.
The differences that caused the most trouble were rarely the visible ones. Silence in a design review means assent in some working cultures and unresolved disagreement in others, and the second kind surfaces later, usually at a point where changing course is expensive. Directness that reads as clarity in one country reads as discourtesy in another. Commitment to a date carries different weight depending on whether the culture treats a plan as a forecast or as a promise. None of this is a matter of one approach being better, and treating it as such is both wrong and, in practice, a fast way to lose good people.
The mechanisms that helped were mostly boring. Writing things down travels across time zones and second languages far better than meetings do, and a written decision with an owner and a date survives contact with a distributed team in a way that a shared understanding reached verbally does not. Most colleagues in a DACH-centred organisation are working in their second or third language, which makes precision and plainness in written communication a leadership responsibility rather than a stylistic preference; idiom, in particular, is a needless source of ambiguity. Making escalation procedural rather than personal lowers the cultural cost of using it, which matters a great deal in environments where contradicting a senior colleague in an open meeting is not something people do lightly. Where junior colleagues were unlikely to disagree publicly, the answer was not to ask them to behave differently but to create a channel where the disagreement could reach me.
Distributed delivery has its own trap. A nearshore or offshore team treated as a capacity pool receiving tickets will never develop judgement, because judgement develops through owning an outcome and living with the consequences. The teams that became genuinely strong were the ones given end-to-end ownership of something a client cared about, with the same standards and the same access to context as anyone else.
The uncomfortable observation, offered with reasonable confidence after several years of this, is that most of the failures I have seen attributed to cultural difference were on closer inspection failures to write down what had been decided.
7. Measure organisational value
An organisation that cannot describe its own performance will eventually be described by whoever has the most convenient numbers. The measures below are the ones that proved informative; every one of them can be gamed individually, which is precisely why they have to be read together.
- Growth and shape. Headcount alone says little. The ratio of senior to junior, and the rate at which juniors become independent, says considerably more about whether the organisation is developing capability or merely absorbing it.
- Utilisation. Necessary to watch and dangerous to maximise. Utilisation pushed to its ceiling consumes exactly the time in which standards, mentoring, and reusable assets are produced, and the cost of that arrives two quarters later in quality and attrition.
- Retention, separated into regretted and non-regretted. An aggregate attrition figure hides the only version of the question that matters.
- Delivery predictability. The variance between what was committed and what was delivered is more diagnostic than throughput, particularly with clients, for whom a predictable team is worth more than a fast one.
- Quality. Defects reaching production, rework as a share of effort, and incident frequency and duration. Rework is the most honest of these and the least frequently tracked.
- Promotion and succession. The proportion of lead and architect roles filled internally is a direct readout on whether the development pipeline is real.
- Client outcomes. Renewal and expansion, and a softer signal that proved surprisingly reliable—whether clients start bringing the organisation their harder problems rather than their better-specified ones.
- Capability maturity. What the organisation can now do repeatably that it could not do a year ago, stated concretely enough that someone could disagree with it.
The value in the set is relational. Utilisation improving while retention degrades is not a good quarter. Delivery predictability improving while rework climbs usually means teams are protecting dates by deferring quality, which is a loan rather than an achievement. I over-indexed on utilisation early, for the entirely understandable reason that it was the number the commercial model was most sensitive to, and it took longer than it should have to recognise what it was quietly consuming.
Closing
None of this resolves into a method. The seven principles interact, they are learned out of order, and each of them was in some way arrived at by first doing the simpler thing and watching it fail at a larger scale. What they have in common is that they are all concerned with reducing an organisation's dependence on particular individuals, including—and perhaps especially—on the person leading it.
A successful Data & AI organisation is not defined by how many models, pipelines, or dashboards it produces. It is defined by whether it repeatedly converts business questions into trusted, governed, and measurable outcomes, without depending on individual heroes.
References
- NIST AI Risk Management Framework: Core — A practical governance reference for explicit roles, accountability, continuous measurement, and risk management across the AI lifecycle.
- Google re:Work: Understand team effectiveness — Research-based guidance on psychological safety, dependability, structure and clarity, meaning, and impact in effective teams.
- DORA: A history of DORA's software delivery metrics — A measurement reference for balancing delivery throughput, stability, and operational performance rather than relying on one convenient metric.
- NeurIPS: Hidden Technical Debt in Machine Learning Systems — Foundational analysis of the operational and organisational complexity surrounding production machine-learning systems.
- Basel Committee: Principles for effective risk data aggregation and risk reporting — Regulated-industry context for governance, ownership, data architecture, accuracy, and timely evidence production.
Author
Géza Kuti is a senior Data and AI executive based in Bülach (ZH), Switzerland, focused on data strategy, enterprise architecture, AI governance, hybrid cloud, and regulated delivery.
Related expertise
Related articles
What AI Did Not Make Cheaper
AI has made software generation cheaper, but verification, evidence, security, accountability, and regulatory defensibility remain the real enterprise cost.
Leading AI Transformation Without Leaving People Behind
AI transformation creates durable value when organisations redesign workflows, roles, learning, accountability, disclosure, and human oversight—not when they merely increase licence utilisation and prompt counts.
When the Model Became a Critical Dependency
The Fable 5 shutdown showed that frontier models can become externally controlled dependencies for regulated enterprises. That changes how firms should think about sovereignty, portability, concentration risk, and operational resilience.