Ronghao Ni

I am a Ph.D. candidate in Electrical & Computer Engineering at Carnegie Mellon University, advised by Prof. Limin Jia and Prof. Lujo Bauer.

My research sits at the intersection of software security, program analysis, and machine learning. I build practical techniques that help developers find, prioritize, and validate vulnerabilities in real-world software ecosystems.

I am actively looking for industry internship opportunities in software security, program analysis, and AI for code. Let’s connect.

Research

I am interested in security techniques that remain useful beyond the lab: tools that can reason about dynamic code, scale to software supply chains, and turn noisy analysis results into evidence a security engineer can act on.

Publications

Conference ranks use the ICORE 2026 conference rankings. Workshop entries state the rank of their main conference; preprints have no rank.

Learning to Triage Vulnerability Reports from Program Analysis: An Empirical Study in Node.js

Ronghao Ni*, Aidan Z. H. Yang*, Min-Chien Hsu, Nuno Sabino, Limin Jia, Ruben Martins, Darion Cassel, and Kevin Cheang

* Equal contribution.

ASE 2026 — ICORE A* — 2026

TL;DR: Machine learning can triage noisy Node.js vulnerability reports while retaining the findings that matter: at 90% recall, the best model removes 75% of benign reports from manual review.

Abstract

Program analysis tools often produce large volumes of candidate vulnerability reports that require costly manual review, creating a practical challenge: how can security analysts prioritize the reports most likely to be true vulnerabilities? This paper investigates whether machine learning can be applied to prioritizing vulnerabilities reported by program analysis tools. We focus on Node.js packages and collect a benchmark of 1,883 Node.js packages, each containing one reported ACE or ACI vulnerability. We evaluate a variety of machine learning approaches, including classical models, graph neural networks (GNNs), large language models (LLMs), and hybrid models that combine GNNs and LLMs, trained on data derived from program analysis tool outputs (NodeMedic-FINE in our case study). The top LLM achieves F1 = 0.915, while the best provenance-graph-based method achieves F1 = 0.904. On the reports that the upstream tool flags but cannot automatically confirm, at a target of recovering 90% of the exploitable reports, the leading model eliminates 75% of the benign reports from manual review. If the best model is tuned to operate at a precision level of 0.8 (i.e., allowing 20% false positives among all warnings), our approach can report 99.2% of exploitable taint flows while missing only 0.8%, demonstrating strong potential for real-world vulnerability triage.

Paper | ASE page

Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning

Ronghao Ni, Mihai Christodorescu, and Limin Jia

Preprint — 2026

TL;DR: LLMVD.js uses autonomous taint reasoning plus exploit confirmation to detect substantially more Node.js package vulnerabilities than prior tools, while producing evidence that can be checked.

Abstract

The rapidly evolving Node.js ecosystem currently includes millions of packages and is a critical part of modern software supply chains, making vulnerability detection of Node.js packages increasingly important. However, traditional program analysis struggles in this setting because of dynamic JavaScript features and the large number of package dependencies. Recent advances in large language models (LLMs) and the emerging paradigm of LLM-based agents offer an alternative to handcrafted program models. This raises the question of whether an LLM-centric, tool-augmented approach can effectively detect and confirm taint-style vulnerabilities (e.g., arbitrary command injection) in Node.js packages. We implement LLMVD.js, a multi-stage agent pipeline to scan code, propose vulnerabilities, generate proof-of-concept exploits, and validate them through lightweight execution oracles; and systematically evaluate its effectiveness in taint-style vulnerability detection and confirmation in Node.js packages without dedicated static/dynamic analysis engines for path derivation. For packages from public benchmarks, LLMVD.js confirms 84% of the vulnerabilities, compared to less than 22% for prior program analysis tools. It also outperforms a prior LLM-program-analysis hybrid approach while requiring neither vulnerability annotations nor prior vulnerability reports. When evaluated on a set of 260 recently released packages (without vulnerability groundtruth information), traditional tools produce validated exploits for few (≤ 2) packages, while LLMVD.js generates validated exploits for 36 packages.

Paper

On the Difficulty of Selecting Few-Shot Examples for Effective LLM-based Vulnerability Detection

Md Abdul Hannan*, Ronghao Ni*, Chi Zhang, Limin Jia, Ravi Mangal, and Corina S. Pasareanu

* Equal contribution. The rank badge refers to the main NDSS conference.

LAST-X @ NDSS 2026 — Main conf. ICORE A* — 2026

TL;DR: Choosing semantically nearby examples can improve few-shot vulnerability detection for Python and JavaScript, but the benefit is inconsistent for C and C++.

Abstract

Large language models (LLMs) have demonstrated impressive capabilities across a wide range of coding tasks, including summarization, translation, completion, and code generation. Despite these advances, detecting code vulnerabilities remains a challenging problem for LLMs. In-context learning (ICL) has emerged as an effective mechanism for improving model performance by providing a small number of labeled examples within the prompt. Prior work has shown, however, that the effectiveness of ICL depends critically on how these few-shot examples are selected. In this paper, we study two intuitive criteria for selecting few-shot examples for ICL in the context of code vulnerability detection. The first criterion leverages model behavior by prioritizing samples on which the LLM consistently makes mistakes, motivated by the intuition that such samples can expose and correct systematic model weaknesses. The second criterion selects examples based on semantic similarity to the query program, using k-nearest-neighbor retrieval to identify relevant contexts. We conduct extensive evaluations using open-source LLMs and datasets spanning multiple programming languages. Our results show that for Python and JavaScript, careful selection of few-shot examples can lead to measurable performance improvements in vulnerability detection. In contrast, for C and C++ programs, few-shot example selection has limited impact, suggesting that more powerful but also more expensive approaches, such as re-training or fine-tuning, may be required to substantially improve model performance.

Paper | arXiv | Workshop page

Mixture-of-Linear-Experts for Long-term Time Series Forecasting

Ronghao Ni, Zinan Lin, Shuaiqi Wang, and Giulia Fanti

AISTATS 2024 — ICORE A — 2024

TL;DR: MoLE lets simple forecasting models specialize by temporal pattern, reducing their error in more than 78% of evaluated settings and reaching state of the art in 68% of reported comparisons.

Abstract

Long-term time series forecasting (LTSF) aims to predict future values of a time series given the past values. The current state-of-the-art (SOTA) on this problem is attained in some cases by linear-centric models, which primarily feature a linear mapping layer. However, due to their inherent simplicity, they are not able to adapt their prediction rules to periodic changes in time series patterns. To address this challenge, we propose a Mixture-of-Experts-style augmentation for linear-centric models and propose Mixture-of-Linear-Experts (MoLE). Instead of training a single model, MoLE trains multiple linear-centric models (i.e., experts) and a router model that weighs and mixes their outputs. While the entire framework is trained end-to-end, each expert learns to specialize in a specific temporal pattern, and the router model learns to compose the experts adaptively. Experiments show that MoLE reduces forecasting error of linear-centric models, including DLinear, RLinear, and RMLP, in over 78% of the datasets and settings we evaluated. By using MoLE existing linear-centric models can achieve SOTA LTSF results in 68% of the experiments that PatchTST reports and we compare to, whereas existing single-head linear-centric models achieve SOTA results in only 25% of cases.

Paper | arXiv | Code

Other Publications (2)

CFT-Forensics: High-Performance Byzantine Accountability for Crash Fault Tolerant Protocols

Weizhao Tang, Peiyao Sheng, Ronghao Ni, Pronoy Roy, Xuechao Wang, Giulia Fanti, and Pramod Viswanath

AFT 2024 — ICORE B — 2024

TL;DR: CFT-Forensics adds cryptographic accountability to crash-fault-tolerant protocols with low overhead, identifying Byzantine replicas when safety is violated.

Abstract

Crash fault tolerant (CFT) consensus algorithms are commonly used in scenarios where system components are trusted - e.g., enterprise settings and government infrastructure. However, CFT consensus can be broken by even a single corrupt node. A desirable property in the face of such potential Byzantine faults is accountability: if a corrupt node breaks the protocol and affects consensus safety, it should be possible to identify the culpable components with cryptographic integrity from the node states. Today, the best-known protocol for providing accountability to CFT protocols is called PeerReview; it essentially records a signed transcript of all messages sent during the CFT protocol. Because PeerReview is agnostic to the underlying CFT protocol, it incurs high communication and storage overhead. We propose CFT-Forensics, an accountability framework for CFT protocols. We show that for a special family of forensics-compliant CFT protocols (which includes widely-used CFT protocols like Raft and multi-Paxos), CFT-Forensics gives provable accountability guarantees. Under realistic deployment settings, we show theoretically that CFT-Forensics operates at a fraction of the cost of PeerReview. We subsequently instantiate CFT-Forensics for Raft, and implement Raft-Forensics as an extension to the popular nuRaft library. In extensive experiments, we demonstrate that Raft-Forensics adds low overhead to vanilla Raft. With 256 byte messages, Raft-Forensics achieves a peak throughput 87.8% of vanilla Raft at 46% higher latency (+44 ms). We finally integrate Raft-Forensics into the open-source central bank digital currency OpenCBDC, and show that in wide-area network experiments, Raft-Forensics achieves 97.8% of the throughput of Raft, with 14.5% higher latency (+326 ms).

Paper | arXiv

Computer-Aided Clinical Skin Disease Diagnosis Using CNN and Object Detection Models

Xin He, Shihao Wang, Shaohuai Shi, Zhenheng Tang, Yuxin Wang, Zhihao Zhao, Jing Dai, Ronghao Ni, Xiaofeng Zhang, Xiaoming Liu, Zhili Wu, Wu Yu, and Xiaowen Chu

Workshop paper; the rank badge refers to the main IEEE BigData conference.

KDDBHI @ IEEE BigData 2019 — Main conf. ICORE B — 2019

TL;DR: A benchmark of modern CNNs, ensembles, and object detectors on two new skin-disease datasets shows both the promise and the difficulty of image-based clinical decision support.

Abstract

Skin disease is one of the most common types of human diseases, which may happen to everyone regardless of age, gender or race. Due to the high visual diversity, human diagnosis highly relies on personal experience; and there is a serious shortage of experienced dermatologists in many countries. To alleviate this problem, computer-aided diagnosis with state-of-the-art (SOTA) machine learning techniques would be a promising solution. In this paper, we aim at understanding the performance of convolutional neural network (CNN) based approaches. We first build two versions of skin disease datasets from Internet images: (a) Skin-10, which contains 10 common classes of skin disease with a total of 10,218 images; (b) Skin-100, which is a larger dataset that consists of 19,807 images of 100 skin disease classes. Based on these datasets, we benchmark several SOTA CNN models and show that the accuracy of skin-100 is much lower than the accuracy of skin-10. We then implement an ensemble method based on several CNN models and achieve the best accuracy of 79.01% for Skin-10 and 53.54% for Skin-100. We also present an object detection based approach by introducing bounding boxes into the Skin-10 dataset. Our results show that object detection can help improve the accuracy of some skin disease classes.

Paper | IEEE

Education