Data Has Its Own Ontology
Abstract
Ontology has moved to the center of AI engineering and data architecture. Organizations increasingly encode customers, accounts, products, contracts, events, actions, and business rules as machine-readable objects and relationships. Semantic-layer architecture likewise places shared business logic between data stores and consuming applications, including AI agents. These developments make machine reasoning over governed organizational knowledge a practical engineering concern. A prior question remains open: what is data? In this paper, ontology means world knowledge and its governing laws, codified into semantic contracts. The word world is used jurisdictionally. A world is a region of objects and questions over which some coherent body of law has governing authority. This definition creates the burden of proof for the paper: analytical data deserve their own ontology only if they exhibit their own objects, conditions of existence and identity, and laws governing a recurring class of questions. The relational model already supplies a powerful ontology of data form. Relations, tuples, attributes, domains, keys, and relational operations give logical structure to data independently of physical storage. In practice, these forms have also served as proxies for data itself. A row stands in for a point, a column for a measure, a GROUP BY for analytical location, and a JOIN for an analytical relation. These proxies often work. Their limits appear when the same table or relational operation can support several analytically different meanings. The Theory of Data begins from another answer. Data is something about something else. A datum is a typed value at a typed analytical point. From that starting point follow universe and existence law, anchor, measure family, sufficient state, lineage, and lawful transformation. These are forms of data-world knowledge whose authority comes from analytical-data law. Four demonstrations make the jurisdiction visible. Arithmetic can compute an average of averages while remaining silent on whether the displayed averages retain sufficient state for exact continuation. A business ontology can correctly state that Products belong to Categories while leaving open whether the relation induces an analytical partition or multiplies contribution. SQL can sum Inventory across time while the resulting quantity ceases to be Inventory. A missing row can coexist with an analytically existing point. In each case, neighboring systems can represent and execute the objects correctly while their native criteria remain insufficient to adjudicate the analytical question. For AI agents, the distinction is practical. Business ontology resolves domain meaning. Relational and schema structure carry a proxy ontology of data form. Semantic-layer technology exposes governed contracts. The Theory of Data adjudicates analytical identity and lawful transformation. The Statistical Bridge governs the later passage from analytical data to statistical evidence and claims. An agent can succeed in the first three layers and still violate analytical-data law. Reliable analytical agents therefore need these jurisdictions to remain explicit.
// Source
Authors: Huayin Wang
Institutions: Open Source Science Project