Research · Online version
The Data Operating System
What DataHub gives you on day one — and the in-house effort to match it
Investment firms often budget a warehouse or lakehouse and assume the operating layer comes with it. It does not. This page lists the capabilities Fencore DataHub ships out of the box, grouped the way operators and configurators use them, with indicative happy-path build effort to handcraft equivalents in-house. Those figures assume you already know what to build; gathering buy-side operating requirements — and changing them after first use — often roughly doubles time to production.
Lakehouses move data; they do not operate it
Microsoft Fabric, Snowflake, and Databricks are excellent at storage, compute, and increasingly at transformation. They do not give you a data operating system: approvals desks, alert triage, config promotion between environments, pipeline designers, and governance that follows the business — role-based access to data in a UI operators can use, and lineage that tracks values through transforms, not just tables.
Fencore DataHub is that operating layer — a no-code data operating system for investment managers. FenWarehouse (Snowflake by default) is the analytical plane fed from the operational store; it is not a substitute for the approvals desk, alert workflows, or pipeline designers that sit above it.
When you compare build vs buy, compare the operating features — not just the warehouse.
How to read the estimates
The tables below are happy-path build time: calendar weeks for a small in-house team (roughly two engineers plus domain input), with AI coding tools, once someone already knows what “good” looks like and the requirement is stable. They are not demo prototypes, and they are not a full programme quote from a systems integrator.
They do not include discovering buy-side operating requirements — four-eye approvals, lineage from pipelines through to alerts and data dashboards, rollback, ownership, exception routing — or the rewrite after operations first use the system. AI speeds implementation of a specified product; it does not invent that specification. Knowing how to prompt an assistant to create end-to-end lineage across pipelines and into alerts, master data, and reporting surfaces is itself a large ask.
- Happy path = coding against a known, stable spec. Low end of a range is closer to a thin MVP; high end is nearer regulated ops parity for that capability.
- Not in the tables = requirements gathering, security and audit design, and scope change after first use. In practice that overhead often roughly doubles calendar time to a used-in-production system.
- AI shortens boilerplate; it does not remove domain design, edge cases, or the need to already understand the operating model you are asking it to build.
- Buying or building a data operating system is not the end of the work: teams still configure pipelines, dictionaries, and rules. Day one means the platform exists — not that every feed is already mapped.
- Fabric, Snowflake, and Databricks still require you to build the product layer on top — connectors alone are not a data operating system.
Read 10–15 months for the ops layer as engineering once the requirement is known. Time to something operations will actually run is often closer to 20–30 months.
For why in-house programmes still miss even those calendars — scope creep was the single most common killer — see The Build Trap at /resources/the-build-trap/full.
Ten capabilities teams underestimate
These are the outcomes DataHub clients use from early go-live. Each would need to be hand-crafted on an in-house or lakehouse-only route.
Multi-eye approvals
Govern business-critical data changes with assignees, history, impacted records, and permission hooks — not a spreadsheet sign-off. Happy-path in-house parity: about 7–9 weeks (thin MVP about 2–3 weeks).
Central data alerts desk
One place to triage quality and process failures, with ownership, custom dashboards, and pipeline linkage. Happy-path parity: about 4–7 weeks (basic list about 1–2 weeks).
Configuration deployment and rollback
Package configuration with dependencies, promote Dev → UAT → Prod, label releases, and roll back safely. Often the most underpriced line item: about 9–15 weeks happy-path for usable parity; rollback alone about 3–5 weeks.
Permissions and row-level access
Users, groups, roles, field filters, and approval controls in one product. Happy-path parity: about 8–12 weeks.
Data lineage
Impact analysis and audit narrative from source through validations and transforms, including into alerts and data dashboards. Happy-path parity: about 7–11 weeks — and only if someone already knows how lineage should stitch across the estate. Prompting AI to invent that end-to-end model is a large ask on its own.
Visual pipeline platform
FenDataFlow, FenMaster, FenRecon, FenRule, and FenSolution designers with governed runners — not scripts in a repo. Credible happy-path suite: about 4–9 months.
Auto-generated master UIs
Explore and maintain golden records from the data dictionary without a separate app per entity. Happy-path parity: about 7–13 weeks.
Events and automation
File, schedule, queue, and data-driven triggers without custom schedulers. Happy-path parity: about 7–9 weeks.
Reconciliation reporting
Business-readable breaks from FenRecon with drill-down and export. Happy-path parity: about 4–8 weeks.
Cori and natural-language operations
Query documentation, configuration, metadata, data, and alerts; build pipelines from plain language — inside the same access controls as the platform. A grounded, governed assistant is months beyond a generic chatbot, and still depends on a specified operating model to query.
Configuration platform (builders)
These are the areas configurators and data engineers use to design how data flows and how the platform behaves.
| Area | What you get on day one with DataHub | Indicative IHB effort |
|---|---|---|
| Data pipelines | Visual ingest → validate → master → reconcile → orchestrate with audit and preview | 4–9 months for a credible suite |
| Data dictionaries | Shared schema layer; masters and UIs stay consistent | 3–7 weeks |
| Global configuration | Sources, targets, connections, formats, failure behaviours, templates | 2–5 months |
| Events and automation | File, queue, schedule, and data triggers with idempotency | 7–9 weeks |
| Deployments | Promote config across environments; history, labels, rollback, restore | 9–15 weeks |
| Report designer | Reports bound to governed data inside the same platform | 4–8 weeks |
| Workspaces | Isolation and packaging for teams or environments | 4–8 weeks |
| Domain applications | Fund hierarchy, performance and attribution, and similar apps on trusted data | 2–5 months per major app |
Operations platform (business users)
These are the desks and tools operations, data stewards, and investment teams use without opening an IDE.
| Area | What you get on day one with DataHub | Indicative IHB effort |
|---|---|---|
| Approvals | Multi-eye control on critical data with audit and impact | 7–9 weeks (MVP 2–3 weeks) |
| Data alerts | Central triage, assignment, resolution, custom views | 4–7 weeks (MVP 1–2 weeks) |
| Reconciliation reports | Readable breaks from FenRecon with traceability | 4–8 weeks |
| Raw data and file upload | Inspect landed snapshots; controlled business-user ingest | 2–4 weeks each |
| Metadata and glossary | Stewardship and ownership on live models | 3–5 weeks |
| Permissions admin | Roles, groups, and data access policies | 8–12 weeks |
| Lineage explorer | Trust and impact before change | 7–11 weeks |
| Dashboard designer and pivot | Self-serve ops and ad-hoc analysis with ACL | 7–9 weeks / 3–5 weeks |
| Tasks and SLA maintenance | Work management and delivery tracking | 3–5 weeks each |
| Operational support console | Jobs, health, audit, messaging — run the platform | 7–9 weeks |
The ops layer alone: bundle math
Teams often believe a lakehouse plus a few scripts covers “data ops.” Stack only the must-have operating slice:
| Bundle | Weeks |
|---|---|
| Approvals + alerts | 11–15 |
| Config promotion + rollback | 9–15 |
| Permissions + audit hooks | 8–12 |
| Lineage + metadata | 9–15 |
| Basic ops UI around pipelines | 7–9 |
| Subtotal (ops + governance layer) | 44–66 weeks ≈ 10–15 months |
Add pipeline designers, dictionaries, events, reconciliation engines, auto master UIs, report designer, and domain applications, and happy-path breadth is easily a year and a half to two and a half years of engineering — before discovery, first-use change, ongoing maintenance, and the scope creep documented in The Build Trap.
What this means for build vs buy
Before you have finished the approvals desk, the alerts desk, and safe promotion of configuration between environments, happy-path engineering is often ten to fifteen months — and you still do not have mastering, recon, or a governed operational store. Once you include requirements gathering and the changes that follow first use, calendar time to a system operations will run is often closer to twenty to thirty months.
That is what Fencore DataHub is designed to deliver on day one: a buy-side data operating system, not another warehouse project. Product detail lives at /products/datahub/full. Financial framing for a business case is at /resources/return-on-investment/full.
If you are budgeting an in-house platform, line-item the operating features — and the cost of learning what those features must do.
Configuration is still work — hard-coded ETL is a different bet
Reaching a DataHub-like platform — whether you buy it or somehow build it — puts you in a place where the estate can be configured by your team, with substantial self-service for operations. Day-to-day triage, approvals, and many pipeline and quality changes need not wait on IT for every adjustment. IT still owns connectivity, security, and harder builds; it does not have to own every mapping change.
That flexibility is not free. Pipelines, data dictionaries, masters, and validation rules still have to be configured. On Fencore DataHub that work is no-code and faster than on most alternatives — but it is still real calendar time. “Day one” means the operating system is available; it does not mean every source and model change is already done.
The other common path is to skip a configuration layer and hard-code ETL: scripts and jobs that move data without a productised designer, promotion model, or ops-facing change path. The first feed can look cheaper and quicker. What you give up is the ability to swap sources, amend the data model, or retarget validations without pulling the IT team into every change. Flexibility moves from the business and data ops back into engineering tickets.
| Path | What you get | What you still pay |
|---|---|---|
| Hard-coded ETL | Faster first feed if the map never changes | Every source swap, model change, or new validation becomes an IT ticket; operations rarely self-serve |
| Configurable platform (DataHub) | Ops can triage, approve, and often adjust without engineering; IT focuses on platform and hard problems | Upfront and ongoing configuration of pipelines, dictionaries, and rules — real work, but not rebuilding designers, lineage, or the alerts desk |
Hard-coding looks cheaper until the first source replacement. A configuration layer looks like extra cost until you count who has to touch every change.
Related resources
- The Build Trap — why 45 investment data leaders would buy instead of build again: /resources/the-build-trap/full
- DataHub full product overview (Cori, operating loop, suite): /products/datahub/full
- Return on investment and business case builder: /resources/return-on-investment/full
- Request a demo: /request-a-demo
Return to the resource summary, or request a demo.
