18 Ways to Verify AI Data Trustworthiness: Ensure Reliable Insights

www.news4hackers.com-18-ways-to-verify-ai-data-trustworthiness-ensure-reliable-insights-18-ways-to-verify-ai-data-trustworthiness-ensure-reliable-insights

18 ways to check whether data can be trusted for AI The European Telecommunications Standards Institute (ETSI) released TR 104 180, a technical document outlining 18 metrics for evaluating data quality to ensure reliability in artificial intelligence applications.

Framework Overview

The framework provides organizations with a structured approach to assess datasets before deployment, offering formulas for each metric. These criteria are organized into four categories. The first group focuses on foundational data attributes, including completeness, accuracy, consistency, and the absence of duplicate entries. The second category evaluates usability factors such as data availability, traceability to original sources, and timeliness. A third set of metrics addresses fairness, examining whether datasets represent diverse populations equitably. The final group prioritizes privacy safeguards, verifying that individual identities remain obscured and sensitive information is adequately protected.

Key Insights from the Report

Diego Lopez, chair of the ETSI Technical Committee DATA, emphasized the importance of quantifiable data quality standards for organizations developing AI systems. He stated that the metrics enable enterprises to determine if their datasets meet critical requirements for trustworthy AI implementations. The report establishes a standardized methodology for data evaluation, supporting more systematic and reproducible assessments as AI technologies advance.

“The metrics enable enterprises to determine if their datasets meet critical requirements for trustworthy AI implementations.”

Validation of the Framework

To validate the framework, researchers applied the 18 metrics to two publicly available datasets. The first dataset comprised sensor data from aircraft engines, which demonstrated strong performance across all criteria. The data was comprehensive, accurate, and maintained stability over time. The second dataset, a U.S. census collection used for income prediction modeling, revealed significant quality concerns.

Dataset Analysis: U.S. Census Data

Bias disparities emerged when analyzing gender representation in the census data. Approximately 31% of male entries were classified as high earners, compared to 11% of female entries. This threefold gap triggered a bias alert under the framework’s fairness metrics. Privacy vulnerabilities also surfaced, with two distinct issues identified.

  • First, combining four attributes—age, race, sex, and country—enabled the identification of specific individuals within the dataset. Some records contained unique combinations that directly linked to personal identities, posing a substantial risk.
  • Second, sensitive fields such as addresses or social security numbers were stored in unencrypted, plaintext format, lacking masking or encryption protections.

Methodology and Collaboration

The TR 104 180 methodology integrates privacy and confidentiality alongside traditional quality measures, applying mathematical formulas to assess compliance. The technical report was developed through collaboration with academic and industry partners, including Sejong University, EGM, TTA, Daejeon University, and CNIT. A corresponding open-source tool was created to automate dataset scoring against the 18 metrics.

The framework addresses growing challenges in AI governance, providing a technical foundation for verifying data integrity. Its application highlights the necessity of rigorous evaluation processes to mitigate risks associated with biased or insecure datasets.



About Author

en_USEnglish