Fali Wang
👋 Hi, I am Fali Wang
I am a last-year Ph.D. candidate in the College of Information Sciences and Technology at The Pennsylvania State University, advised by Prof. Suhang Wang in the Data Science and Machine Learning Lab.
I have interned at Microsoft (Redmond), Amazon (Palo Alto), and NEC Laboratories America (Princeton), working on test-time scaling, efficient AI, and knowledge-enhanced LLMs.
fqw5095 [at] psu [dot] edu for opportunities or research collaboration.Research Interests
I work on efficient and trustworthy AI: allocating models and compute to what each task actually needs, making language models and graph learners robust and reliable, and studying how graphs and LLMs can enhance each other. My work spans several connected directions:
- Efficient AI & test-time scaling. Compute-optimal budget allocation at the task and query level, adaptive model routing, and agents that search scaling strategies for complex multi-stage tasks (AgentTTS, AgentRE).
- Small language models. SLMs as efficient foundations for next-generation AI: capability enhancement under limited compute, SLM–LLM collaboration, and cloud-edge deployment with privacy and trustworthiness (SLM Survey, SLM–LLM Collaboration).
- Agent-in-the-loop & self-improving AI. Agents that iteratively optimize their own workflows — collaboration topology, model and role assignment, memory, and compute — by accumulating and reusing knowledge from prior trajectories and feedback (AgentTTS, AgentRE, DC-GST, HC-GST).
- Graph learning & graphs for LLMs. Graph self-training under distribution shift, LLM graph reasoning and its benchmarking, knowledge-graph-enhanced LLMs, and graph-enhanced retrieval-augmented generation (DC-GST, HC-GST, GraphSkill, Graphs for LLMs, InfuserKI).
- Trustworthy AI. Robustness, security, privacy, and reliability across language models and graph learning: distribution shift and robustness in GNNs, certified robustness of BERT, backdoor attacks and defenses, machine unlearning, privacy risks in vision-language models, hallucination mitigation, and vulnerabilities in RAG and cloud-edge AI systems (MacroBERTDC-GST, HC-GST, InfuserKI, SLM–LLM Collaboration).
News
| 09/2026 | AgentRE, which generalizes test-time compute-optimal scaling as an optimizable graph, is accepted to NeurIPS 2026. |
| 05/2026 | Started as an Applied Scientist Intern at Microsoft, Redmond. |
| 02/2026 | Graphs for LLMs, our survey on graph-assisted large language models, is accepted to ACL 2026 (Findings). [GitHub] |
| 09/2025 | AgentTTS, an LLM agent for test-time compute-optimal budget allocation, is accepted to NeurIPS 2025. |
| 08/2025 | Our Small Language Models survey is accepted to ACM TIST. |
| 08/2025 | Organized the KDD 2025 Tutorial on Small Language Models and the KDD 2025 Workshop on LLMs for E-Commerce. |
| 04/2025 | Invited talk on SLMs at the WWW 2025 LLM for E-Commerce Workshop. [Slides] |
| 01/2025 | Invited talk on SLMs at Amazon. [Slides] |
Earlier news
| 12/2024 | Started as an Applied Scientist Intern at Amazon, Palo Alto. |
| 11/2024 | Led and released the Small Language Models survey. [arXiv] [GitHub] |
| 11/2024 | Passed the comprehensive exam. |
| 09/2023 | Visiting research intern at NEC Laboratories America, Princeton. |
| 05/2023 | Passed the qualifying exam and became a Ph.D. candidate. |
| 08/2022 | Began my Ph.D. at Penn State University. |
Selected Publications
First-author and co-first-author work, most recent first. * denotes equal contribution. Full list on Google Scholar.
Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph
NeurIPS 2026AgentRE recasts test-time compute-optimal scaling as search over an optimizable graph of models, roles, and budgets, letting an agent jointly decide what to run and how much compute to spend.
GraphSkill: Documentation-Guided Hierarchical Retrieval-Augmented Coding for Complex Graph Reasoning
KDD 2026Lets LLMs solve complex graph-reasoning problems by hierarchically retrieving graph-library documentation and writing executable code, instead of reasoning over graphs in text.
Graphs for LLMs: A Survey of Graph-Assisted Large Language Models
ACL 2026 FindingsA systematic survey of how graph structures assist LLMs — in retrieval, reasoning, planning, agents, and evaluation — with an open paper collection.
AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex Tasks
NeurIPS 2025An LLM agent that iteratively searches the compute-optimal test-time scaling strategy for multi-stage complex tasks — which model to use and how much compute to allocate to each subtask.
The first comprehensive survey of SLMs: architectures, training and enhancement techniques, applications, SLM–LLM collaboration, and trustworthiness. Presented as a lecture-style tutorial at KDD 2025 and in invited talks at Amazon and WWW 2025.
Organizes SLM–LLM collaboration patterns by the goals they serve — performance, cost, cloud-edge privacy, and trustworthiness — and maps open problems.
Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models
Preprint 2026A unified benchmark that evaluates LLM graph reasoning along multiple dimensions of complexity, revealing where current models break down.
Integrates new knowledge-graph facts into an LLM through an infuser that selectively injects only what the model does not already know, reducing hallucination without forgetting.
HC-GST: Heterophily-aware Distribution Consistency-based Graph Self-training
CIKM 2024Graph self-training that selects pseudo-labels to keep the homophily distribution of the training set consistent with the full graph, so heterophilic nodes are no longer under-represented.
Distribution Consistency-based Self-Training for Graph Neural Networks with Sparse Labels
WSDM 2024Chooses pseudo-labeled nodes that shrink the distribution gap between labeled and unlabeled nodes, making GNN self-training reliable when labels are scarce.
More publications & collaborations
- Retrieved But Not Reliable: A Survey on Attacks, and Defenses in Retrieval-Augmented Generation.
EMNLP 2026 - Adversarial Reinforcement Learning for Robust Diffusion Large Language Model Unlearning.
ICML 2026 - Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation.
ICLR 2026 - Bradley-Terry and Multi-Objective Reward Modeling Are Complementary.
ICLR 2026 - How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use.
ICLR 2026 - Image Corruption-Inspired Membership Inference Attacks against Large Vision-Language Models.
EACL 2026 - BioMol-MQA: A Multi-Modal Question Answering Dataset for LLM Reasoning over Bio-Molecular Interactions.
ICDM 2026 - Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking.
NeurIPS 2025 - SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.
TMLR 2025 - Catastrophic Failure of LLM Unlearning via Quantization.
ICLR 2025 - Enhance Graph Alignment for Large Language Models.
Neural Networks - Maximum Entropy Loss, the Silver Bullet Targeting Backdoor Attacks in Pre-trained Language Models.
ACL 2023 Findings - Dynamic Graphs and Large Language Models: A Survey of Mutual Enhancement.
Preprint 2026 - Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration?
Preprint 2026 - MacroBERT: Maximizing Certified Region of BERT to Adversarial Word Substitutions.
DASFAA 2021 - ConvMB: Improving Convolution-Based Knowledge Graph Embeddings by Adopting Multi-Branch 3D Convolution Filters.
ISPA 2021 - BEFSR: A Multiple Attention-Based Model Considering Bidirectional Entity Information Flows and Few-shot Relations.
ICPR 2022 - De-Co: A Two-Step Spelling Correction Model for Combating Adversarial Typos.
ISPA 2020 - NarGNN: Narrative Graph Neural Networks for New Script Event Prediction Problem.
ISPA 2020 - Research and Simulation on Processing Speed Connection of Multi-axis Woodworking Engraving Machine.
Journal of Northeast Forestry University 2018
Industry Experience

Applied Scientist Intern · Redmond, WA · Mentors: Dr. Zhenwei Dai, Joy Zeng, and Dr. Qi He
Query-level compute-optimal budget allocation for efficient LLM test-time scaling.


Research Intern · Princeton, NJ · Mentors: Dr. Runxue Bao and Dr. Haifeng Chen
Knowledge infusion to mitigate hallucination in LLMs → InfuserKI (EMNLP 2024).
Education

Ph.D. in Informatics, College of Information Sciences and Technology · Advisor: Prof. Suhang Wang

M.Eng. in Software Engineering, School of Cyber Security

B.Eng. in Software Engineering · Ranked 1st in the cohort
Academic Activities & Service
Organizer

A Tutorial on Small Language Models in the Era of Large Language Models: Architecture, Capabilities, and Trustworthiness

Invited Talks
- Small Language Models in the Era of LLMs · WWW 2025 Workshop on LLMs for E-Commerce · 04/2025 [Slides]
- Small Language Models in the Era of LLMs · Amazon · 01/2025 [Slides]
Service
- Web Chair · KDD 2027
- Guest Editor · ACM Transactions on Intelligent Systems and Technology (TIST)
- Program Committee · ICMR 2026, IEEE BigData 2026, ACM MM 2026
- Reviewer · ICLR, ICML, NeurIPS, KDD, ACL, EMNLP, WWW, CIKM, IJCAI, SDM, RecSys, ACM MM, IEEE BigData; ACM TIST, ACM Computing Surveys
- Volunteer · NeurIPS 2025
Teaching
- Teaching Assistant · DS 305: Algorithmics · Penn State · Spring & Fall 2026
- Teaching Assistant · DS 420: Network Analytics · Penn State · Fall 2024
Honors & Awards
- Travel Award, PSU College of IST · CIKM 2024 & EMNLP 2024
- National Scholarship, Ministry of Education of China · 2015 & 2016
- Honorable Mention, COMAP Interdisciplinary Contest in Modeling · 2018
- First Prize (Provincial) & Third Prize (National), Lan Qiao International Programming Contest · 2017
- Top 10 Media Person, China College Students Online Campus Netcom, Ministry of Education · 2017
- Second Prize, CSIAM National Undergraduate Mathematical Contest in Modeling · 2016
Visitor map










