Evaluating Privacy Compliance with LLMs: a Cross-Jurisdictional Multi-Law Reasoning Framework Based on Contextual Integrity Theory

The Hong Kong University of Science and Technology
Department of Computer Science and Engineering


MPhil Thesis Defence


Title: "Evaluating Privacy Compliance with LLMs: a Cross-Jurisdictional 
Multi-Law Reasoning Framework Based on Contextual Integrity Theory"

By

Miss Zirui WANG


Abstract:

Large Language Models (LLMs) demonstrate high potential for assessing privacy 
compliance. However, current research is limited to the formulation of a 
single jurisdiction and a single regulation, and it does not reflect 
multidimensional cross-statutory reasoning that is common in real judgments. 
In addition, the cross-jurisdictional and cross-lingual transferability of 
LLMs has not been systematically assessed.

To fill this gap, this study systematically evaluates the effectiveness of
LLMs in English-language legal case judgments based on China's four personal
data protection laws: the Personal Information Protection Law (PIPL), the
Cybersecurity Law (CSL), the Data Security Law (DSL), and the Basic Security
Requirements for Generative AI Services (GenAI). Based on the theory of
Contextual Integrity (CI), the study creates a structured knowledge base
covering these four laws and compiles an interjurisdictional assessment
dataset including 1,406 English case profiles. We test three prompt
strategies—Direct Prompting (DP), Automated Chain-of-Thought (CoT-auto), and
Retrieval-Augmented Generation (RAG) implemented via the BM25-content
algorithm—on six LLMs (GPT-4-turbo-04-09, GLM-4-9B-Chat, Llama-3-8B-Instruct,
Mistral-7B-Instruct-v0.2, Qwen2.5-7B-Instruct, Qwen3-8B).

Empirical research reveals four main findings. First, the optimal setting
(GPT-4-turbo-04-09 + BM25-content) achieves 86.0% accuracy, but because of
compounded difficulties such as terminological mapping, semantic alignment
between languages, and differences in the legal system, a notable gap
persists relative to prior single-statute, intra-lingual HIPAA-based
benchmarks, underscoring the inherent difficulty of multi-statute,
cross-lingual legal reasoning. Second, RAG is a double-edged sword: it
increases the accuracy of GPT-4-turbo-04-09 by 21 percentage points, but
reduces the performance of GLM-4-9B-Chat and Qwen2.5-7B-Instruct, indicating
that RAG efficacy depends heavily on the model's inherent filtering
capability. Third, GenAI is the biggest challenge, with under-selection
reaching 41.9% for GPT-4-turbo-04-09 under BM25-content, reflecting the
difficulties posed by novelty and technology-driven features. Fourth,
Clearview AI's case study further confirms that RAG via the BM25-content
method fully covers the legal requirements (no omissions, zero over-selection
rate), while DP and CoT-auto omit 81.8% and 54.5% of the applicable
provisions, respectively.

This study establishes the first multi-statute collaborative privacy
compliance framework. It systematically investigates the
cross-jurisdictional transferability of LLMs and provides actionable
insights for building credible legal AI systems through rigorous
benchmarking across diverse models and strategies.


Date:                   Tuesday, 29 September 2026

Time:                   9:00am - 11:00am

Venue:                  Room 3494
                        Lifts 25/26

Chairman:               Dr. Dan XU

Committee Members:      Dr. Yangqiu SONG (Supervisor)
                        Prof. Raymond WONG