top of page

Semantic Interoperability Across Data Spaces: IntuView's Ontological Mediation Technology in DS2

  • Laura Gavrilut
  • Jul 13
  • 5 min read

As governments, enterprises, and research organizations generate unprecedented volumes of information, one of the greatest challenges of the digital age is not collecting data, it is finding, understanding, and sharing the right information across languages, cultures, and organizational boundaries.


This challenge lies at the heart of the European vision for interoperable data spaces. While significant progress has been made in data governance, privacy, and information sharing protocols, a critical obstacle remains - enabling users to discover relevant information regardless of the language, terminology, or data model in which that information was originally created. Traditional search technologies often struggle in multilingual and multi-domain environments, resulting in either low recall, where important information is missed, or low precision, where users are overwhelmed by irrelevant results.


To address this challenge, IntuView is contributing its patented semantic intelligence technology to the DS2 Horizon Europe project through the development of the Common Language Mediation (CLM) Module. The CLM serves as a semantic interoperability layer that enables information originating from different data spaces to be represented, searched, and understood through a common conceptual framework.


At the core of IntuView's technology is its Meaning Mining™ approach, implemented through its flagship platform, IntuScan. Unlike conventional search and text analytics systems that rely primarily on keywords, lexical matching, or machine translation, IntuScan converts both unstructured text and semi-structured metadata into ontology-based semantic representations. This process, known as ontologization, transforms information into a language-independent, hierarchical, and unambiguous representation that captures not only the explicit content of a document but also its context, relationships, sentiment, and implicit meaning.


Within DS2, this ontological representation acts as a common semantic "lingua franca" shared across participating data spaces. Information submitted by any participant is analyzed and converted into an ontological digest containing concepts, attributes, relationships, and hierarchical structures. User queries undergo the same transformation process. By matching ontological representations rather than surface-level words, users can search in one language and discover relevant information stored in another, while preserving the original meaning and context of the information.


The hierarchical nature of the ontology significantly enhances information discovery. Unlike traditional search systems that return only exact or similar terms, the CLM module understands conceptual relationships. A search for a broad concept such as "Britain" can automatically retrieve information relating to England, Manchester, or other subordinate concepts. Likewise, equivalent concepts expressed through different terminology, such as "Britain" and "United Kingdom," can be recognized as representing the same underlying meaning. This enables richer, more accurate retrieval of information across domains and languages.


The CLM module performs two principal functions within the DS2 ecosystem. First, it normalizes information stored in the Catalogue by converting metadata and textual content into a common ontological format that bridges semantic differences between data spaces. Second, it supports the Dialogue and Retrieval Component (DARC) chatbot by interpreting natural-language queries, resolving ambiguities, and identifying the intended meaning of user requests.


This capability is particularly important when dealing with ambiguous terminology. For example, a user searching for information about "courts" may be referring to courts of law, sports courts, royal courts, or educational courtyards. Because the CLM lexicon maps lexical expressions to multiple ontological concepts, the system can engage the user in a clarifying dialogue and identify the intended meaning before executing a search. Once disambiguated, the query is matched against similarly normalized information stored in the Catalogue, ensuring accurate and contextually relevant results.


The module incorporates a comprehensive natural language processing pipeline consisting of tokenization, part-of-speech tagging, morphological analysis, feature extraction, regular-expression processing, multilingual lexicons, named-entity recognition, rule-based reasoning, document categorization, and semantic summarization. These components work together to transform language-dependent information into a searchable ontological structure while preserving the relationships and contextual information required for advanced reasoning.


Beyond information retrieval, the CLM module supports multilingual information fusion and semantic knowledge discovery. By representing information from different languages, cultures, and domains within a unified conceptual framework, the system can identify hidden relationships, discover implicit affinities, recognize inconsistencies, and uncover insights that would remain invisible in traditional search environments.


The technology also contributes to data quality and governance through anomaly detection and ontological comprehensiveness control. Metadata originating from different data spaces may contain inconsistencies, unusual values, or previously unseen concepts. The CLM module analyzes incoming information to identify anomalies, detect gaps in existing ontologies, and generate alerts when new concepts emerge that require extension of the semantic model. This capability ensures that the platform evolves alongside the changing data ecosystem.


A key research objective within DS2 is the development of semi-automated ontology generation and extension. Traditionally, building domain ontologies has been a labor-intensive process requiring extensive expert involvement. The project is therefore developing methods that combine machine learning, document clustering, TF-IDF analysis, thesaurus exploration, synonym mapping, concept extraction, and statistical categorization to accelerate ontology creation. These methods enable the platform to identify emerging concepts, propose new ontological structures, and support the continuous expansion of semantic coverage as new domains and data products are introduced.


The CLM module is tightly integrated with other components of the DS2 platform. Ontological digests generated by the module are stored within the Catalogue to support semantic search and discovery. The DARC chatbot uses the module to understand user intent and conduct multilingual dialogue. The module can also contribute to orchestration and governance functions by providing a common semantic interpretation of policies, procedures, and domain-specific information exchanged between participating data spaces. Dedicated REST APIs enable seamless integration with all relevant DS2 components.


Importantly, the CLM module functions as a true Data Space Intermediary. Rather than simply connecting databases, it provides semantic normalization between heterogeneous data spaces, ensuring that information can be understood consistently regardless of language, terminology, or local data conventions. This semantic interoperability enables data spaces to communicate at the level of meaning rather than merely exchanging data.


The significance of this approach extends beyond the immediate objectives of DS2. In the era of Artificial Intelligence, ontology-based semantic representations provide a language-independent foundation for explainable and scalable AI systems. By abstracting meaning from language, organizations can reduce dependence on maintaining separate language-specific models while improving transparency, traceability, and interoperability.


IntuView's semantic intelligence technologies have already been deployed by defense, intelligence, law enforcement, and security organizations across the United States, Europe, and the Middle East. Through the DS2 project, these capabilities are now being applied to one of Europe's most ambitious digital transformation initiatives: enabling seamless, intelligent access to information across interconnected data spaces.


By combining advanced semantic analysis, multilingual understanding, ontology-driven AI, and automated knowledge engineering, IntuView's CLM technology helps create a future in which information can move freely across linguistic, organizational, and technical boundaries allowing users not only to find the information they seek, but also to discover relationships and insights they did not know existed.

Comments


bottom of page