Can LLMs Truly Reason? From Stochastic Parrots to In-Context Learners

The Hong Kong University of Science and Technology
Department of Computer Science and Engineering


PhD Thesis Defence


Title: "Can LLMs Truly Reason? From Stochastic Parrots to In-Context 
Learners"

By

Miss Tsz Ting CHUNG


Abstract:

The rapid advancement of Large Language Models (LLMs) has sparked a 
fundamental debate: are these models genuinely reasoning, or are they merely 
stochastic parrots? This thesis addresses this question through a 
comprehensive evaluation framework organized around four criteria that a 
truly reasoning-capable LLM must satisfy: (1) going beyond memorization, (2) 
closing the human-machine gap, (3) reasoning traceability, and (4) handling 
global dependencies.

Guided by these criteria, we systematically investigate LLM reasoning across 
three complementary domains, each exposing deficiencies along different 
dimensions. In the physical domain, we design paired tasks at two cognitive 
levels to isolate memorization from genuine understanding, revealing a ~40% 
accuracy gap between LLMs and humans and thereby exposing the stochastic 
parrot phenomenon. In the logical domain, we construct counterintuitive 
premises grounded in propositional logic to neutralize commonsense shortcuts, 
showing that standard LLMs achieve near-random performance (<36%) while 
humans reach 86.7%, demonstrating that LLMs lack genuine logical deduction 
when semantic heuristics are stripped away. Extending our investigation to 
the long-context domain, we design a benchmark satisfying also the global 
dependency, where correct answers cannot be obtained through local retrieval 
pattern-matching alone but require reasoning over globally scattered 
evidence. Results reveal that reasoning-oriented LLMs lag behind humans by 
>15% in accuracy and >30% in reasoning quality.

Each evaluation study highlights where LLMs must improve along different 
criteria, collectively painting a comprehensive picture of current reasoning 
limitations. We then shift from diagnosing failures to enabling genuine 
learning by reframing context as structured guidance. For static reasoning 
tasks, we show that many-shot CoT-ICL works best when demonstrations, viewed 
as reasoning trajectories, are understandable and smoothly sequenced as a 
pedagogical curriculum. Extending this trajectory-based view to dynamic 
agentic settings, we further investigate whether insights distilled from past 
trajectories can guide future decisions when retrieved according to the 
agent's current state and next-action bottleneck. This formulation addresses 
a key limitation of existing embedders, which primarily capture surface-level 
semantic similarity rather than progress-oriented action relevance. 
Collectively, this thesis charts a path from stochastic parrots to in-context 
learners guided by curriculum-structured demonstrations and reusable 
insights.


Date:                   Monday, 13 July 2026

Time:                   10:00am - 12:00noon

Venue:                  Room 3494
                        Lifts 25-26

Chairman:               Prof. Mengze SHI (MARK)

Committee Members:      Prof. Dit-Yan YEUNG (Supervisor)
                        Prof. Albert CHUNG
                        Prof. Raymond WONG
                        Dr. Jing TANG (EMIA)
                        Dr. Hongsheng LI (CUHK)