Research
Interpretable imaging algorithms, medical image analysis, and applied machine learning.
Interpretable imaging algorithms, medical image analysis, and applied machine learning.
Peer-reviewed work in image restoration, computational imaging, medical image analysis, and robust learning.
ECCV 2026 · Penn State
Recovers sharp images and unknown blur kernels by learning Fourier phase and amplitude updates; improves deblurring under noise and limited training data.
Problem and method. Recovering a sharp image from unknown blur requires estimating both the image and its blur kernel. In this co-authored work, UPADNet separates Fourier phase and amplitude, derives linear minimum mean-square error (LMMSE) estimators, and learns the updates of an alternating optimization algorithm. Each stage refines image and kernel estimates, with generated weights and multiple scales adapting the reconstruction.
Results. The 27M-parameter model reaches PSNR/SSIM of 34.63 dB/0.975 on GoPro, 41.29 dB/0.979 on RealBlur-R, and 34.35 dB/0.946 on RealBlur-J. Controlled COCO experiments show stronger restoration under added noise and reduced training data; using 60% of the training set, it reaches 32.12 dB/0.926.
Architecture, restoration examples, and robustness experiments.
WACV Workshops 2026 · Dolby Laboratories / Penn State
Recovers sharp images from motion blur by combining event-camera measurements with physical blur modeling and learned image priors.
Problem and method. Long exposures merge fast motion into a blurred frame; event cameras retain the brightness changes within the exposure. I developed EUMD to connect these measurements to a physical blur operator, then unroll half-quadratic splitting into eight reconstruction stages with a Transformer-enhanced U-Net prior. Training-time timestamp correction aligns sharp labels with the event-derived reconstruction; evaluation uses the exposure midpoint.
Results. EUMD reaches 37.35 dB/0.9801 on synthetic GoPro and 32.42 dB/0.937 on real EVRB. PSNR improves by 0.31 dB over ClearSight on GoPro and 1.04 dB over CMTA on EVRB. The Transformer and training timestamp correction contribute 0.41 dB and 1.02 dB, respectively, in the paper’s ablations.
Event-based reconstruction architecture and visual comparisons.
Journal of Neurosurgery: Pediatrics 2026 · Penn State
Uses the unaffected brain hemisphere to assess growth in pediatric hydrocephalus when shunt artifacts prevent reliable whole-brain MRI measurements.
Problem and method. Metallic shunt hardware can obscure part of the brain in MRI. This co-authored study measures the artifact-free hemisphere as a proxy for whole-brain growth. A Dense U-Net segments brain tissue and cerebrospinal fluid (CSF), a second model identifies the hemibrain, and voxel dimensions convert the masks into volumes. The training loss combines Dice overlap with CSF-focused total-variation regularization.
Results and scope. MRI from 75 ESTHI trial patients supports the hemisphere-based approach: postoperative hemisphere ratios stabilize, and hemibrain and whole-brain growth show similar patterns where both are measurable. Brain and CSF segmentation achieve Dice scores of 94.7% and 94.2%. My contributions include data interpretation and statistical analysis. Some scans require manual refinement; validation in larger, more diverse cohorts remains future work.
Volumetry, segmentation, and brain-growth analysis.
Neuroelectronics 2025 · Penn State
Predicts CT-informed phase corrections in 0.45 seconds to improve transcranial ultrasound focusing in acoustic simulations.
Problem and method. The skull distorts focused ultrasound, while simulation-based correction is slow. DB-SIPAC combines a direct-propagation pathway branch with a full-skull branch that captures broader reflections and refractions. Iterative Time Delay Search (ITDS) refines the training targets. My contributions included methodology, software, data preparation, and validation.
Results and scope. The dataset contains 1,260 skull/focal-point pairs from 36 CT-derived slices of 12 adults. In 2D k-Wave simulations, DB-SIPAC predicts delays in 0.45 seconds, versus 32 seconds for time reversal and 249 seconds for hybrid angular spectrum, and reaches mean normalized focal pressure of 0.9923. Reduced-data tests and saliency maps support the anatomical design. These are simulation results; clinical validation remains future work.
Network design, acoustic setup, and simulation results.
IEEE TCI 2024; ICIP 2023 · Penn State
Restores images with known blur kernels using a compact unrolled network that retains provable convergence and performs well with limited training data.
Problem and method. Learning independent parameters at each layer can remove an iterative algorithm’s convergence guarantees. In this co-authored work, DCUNBD (ICIP 2023) and its journal extension DECUN (IEEE TCI 2024) unroll half-quadratic splitting and constrain learned filters and thresholds to approach limiting values. The analysis proves convergence as network depth increases and characterizes its rate.
Results. A 30-layer DECUN configuration uses only 1,117 trainable parameters. Tests with linear and nonlinear blur kernels show competitive restoration quality; experiments using 10% of the training images demonstrate the model’s data efficiency. Convergence curves and image comparisons connect the theoretical guarantees to observed results.
Network design, convergence curves, and restoration results.
NTIRE 2023 · IPAL-Bokeh / Penn State
Transforms out-of-focus lens bokeh while preserving a sharp foreground; team IPAL-Bokeh achieved the lowest real-image LPIPS in NTIRE 2023.
My participation and method. As a member of IPAL-Bokeh, I contributed to the Segmentation-Guided Lens-Mapping Scheme (SGLMS). Detail and semantic branches estimate the foreground mask; lens-specific encoders and decoders transform the background bokeh. Mask-guided fusion preserves the sharp foreground, while gated convolutions and selective-kernel fusion improve the mapping.
Results. Our 7M-parameter submission uses no ensembling and achieves the lowest real-image LPIPS (0.2161) among submitted methods. This distinction measures perceptual similarity on real captures; the overall challenge ranking uses PSNR/SSIM. The challenge report documents our team’s participating method.
Segmentation, lens mapping, and challenge results.
CVPR Workshops / NTIRE 2023 · iPAL-LightDehaze / Penn State
Removes spatially varying haze with a lightweight physical-model and Transformer approach; team iPAL-LightDehaze ranked 4th in LPIPS at NTIRE 2023.
My participation and method. As a member of iPAL-LightDehaze, I contributed to TransER for spatially varying haze. A shared Transformer–convolution encoder and three decoders estimate atmospheric light, transmission, and a haze-free image. A lightweight ensemble network then fuses physical-model and direct reconstructions, using feature distillation from a clean-image reconstruction teacher.
Results. The 2.6M-parameter model reaches PSNR/SSIM of 17.03 dB/0.597 on Dense-Haze and 21.64 dB/0.743 on NH-Haze. Our NTIRE 2023 submission ranks 4th in LPIPS among 17 finalists (0.384), without extra training data. Reported inference takes 0.72 seconds on a Titan XP GPU.
Method paper · Challenge report · Code
Architecture, feature-fusion modules, and benchmark results.
NTIRE 2021 · iPAL-GridFFA / Penn State
Removes non-uniform haze with grid-based feature fusion and attention; team iPAL-GridFFA placed 2nd in SSIM and tied for the lowest reported VGG-based LPIPS at NTIRE 2021.
My participation and method. As a member of iPAL-GridFFA, I contributed to a generative adversarial network (GAN) for non-uniform haze. Its 3 × 6 GridDehazeNet-based generator combines feature-fusion groups, channel and pixel attention, and skip connections. Each group contains 15 basic blocks, while a PatchGAN-style discriminator guides restoration quality.
Results. Our submission ranks 2nd in SSIM (0.839) among 23 finalists and ties for the lowest VGG-based LPIPS (0.194 at the precision reported). These metric-specific results appear in Table 1; the report’s final perceptual ranking is based on mean opinion score.
GridFFA, Grid Net, feature attention, and challenge results.
2018–2021 · Xi’an Jiaotong University
Developed collaborative whole-slide annotation software and renal cancer analysis models, with peer-reviewed studies, AIPath datasets, and a granted annotation patent.
Collaborative annotation. Gigapixel pathology slides are difficult to inspect and annotate consistently. I contributed to OpenHI/OpenHI2 and developed cloud-based collaborative annotation software. Multi-scale superpixels support pixel-level semantic labeling and responsive region retrieval; OpenHI2 adds calibrated magnification, shared diagnostic regions, agreement tests, and modular image analysis. OpenHI reports roughly 300 ms typical response time under its test conditions.
AI-assisted workflows. My software work included AI-assisted labeling and annotation workflows. Related co-authored PIMIP research integrates slide management, multi-device collaboration, automatic nucleus labeling, manual correction, extensible analysis, and linked patient/report information.
Nuclei grading and datasets. I developed renal cell carcinoma (RCC) machine-learning models and co-authored related nuclei-grading studies. CHR-Net uses W-Net to separate crowded nuclei, then combines high-resolution features and two cross-category classification heads. The MICCAI study contains 1,000 patches and 70,945 annotated nuclei, including clear-cell and supplementary papillary RCC samples. CHR-Net reaches Dice 0.8790 and average class-wise panoptic quality 0.5458 in the reported evaluation.
Whole-slide and multi-source analysis. Other co-authored studies combine tumor detection, RCC subtyping, grading, and whole-case summaries using TCGA and hospital slides. A personalized framework links similar histology with clinical reports and survival analysis. An attention-based graph convolutional network (GCN) extracts structured information from 3,632 TCGA reports across four cancer types, improving macro-F1 on most evaluated tasks. Annotation-granularity experiments show how precise labels affect model performance.
Outcomes. This work connects reusable research software, annotated datasets, and image/text analysis. I am a co-inventor on the collaborative annotation patent US11392634B2, granted in 2022.
OpenHI paper · OpenHI2 paper · CHR-Net paper · PIMIP paper · RCC framework · Personalized analysis · Report extraction · Annotation study · OpenHI code · AIPath datasets · ccRCC grading dataset
Platform workflows, annotation tools, and research evaluations.
NLPCC 2021 · Xi’an Jiaotong University
Uses meta-learning to handle noisy labels and class imbalance in literary-review sentiment analysis, improving macro-F1 by 13.51 percentage points.
Problem and method. Star ratings produce noisy sentiment labels, while positive reviews dominate the training data. This co-authored study introduces BERT-MLB: a BERT classifier with a Looking Back meta-model. An LSTM learns each sample’s training weight from historical features and its label embedding, guided by a small, balanced set of manually verified reviews.
Results. The dataset contains 109,286 reviews of 187 Chinese literary books, with 600 verified meta samples and a balanced 6,000-review test set. Macro-F1 rises from 61.12% for BERT to 74.63% for BERT-MLB, a gain of 13.51 percentage points. Sample-weight analysis shows reduced influence from noisy labels and greater attention to minority classes.
Meta-learning architecture and noisy-label analysis.
Preliminary research on Poisson-noise formulations for diffusion-based generation, motivated by signal-dependent imaging noise. Explored connections between Poisson processes and diffusion models, with low-light and night-scene generation as target applications.
TypeScript · React Native · Agentic AI
Built an Android application that converts multi-turn text and image conversations into structured prompts and AI video-generation tasks. Implemented persistent task state, asynchronous job tracking, retries, recovery, and configurable workflows with validation and runtime tests.
LLMs · Multi-Agent Systems
Built a novel-generation workflow with planning, writing, review, and revision agents. Persistent character profiles, world state, and event timelines support narrative continuity; review loops identify and address inconsistencies in long-form output.