Curriculum Vitae
Education
- University of California San Diego, Sept. 2020 - June 2025
- Ph.D in Electrical Engineering, Overall GPA: 4.0/4.0
- Advisor: Prof. Truong Q. Nguyen, VP Lab
- Research: Advancing Automated AMD Stage Grading with 2D and 3D Retinal OCTA Imaging (link)
- University of Science and Technology of China, Sept. 2017 - July 2020
- Master in Information and Communication Engineering, Overall GPA: 4.01/4.3
- Advisor: Prof. Dong Liu, VIDAR Group
- Research: Action recognition oriented video super-resolution
- University of Science and Technology of China, Aug. 2013 - June 2017
- B.S. in Electronic and Information Engineering, Overall GPA: 3.73/4.3
Work Experience
- Applied Scientist, July 2025 – Present
- AGI, Amazon, Sunnyvale
- Contributed to multimodal LLM and image/video generative model development through large-scale data curation, distributed post-training, and standardized benchmark evaluation.
- Data:
- Designed and delivered multi-million–scale PT/SFT datasets for Nova 2.0 (VLLM), including 1.36M logo videos, 377K celebrity localization samples, and a 200M knowledge-enriched concept dataset (logos, landmarks, animals, plants, fungi).
- Developed scalable multi-stage caption rewriting pipelines.
- Built a video semantic taxonomy framework to help content-aware curation for targeted semantic improvements.
- Model:
- Conducted fine-tuning and alignment experiments on 32–64 GPU nodes using PyTorch, validating new data through ablations to prevent regressions.
- Trained image generation models at scale on 32 nodes for data selection ablations.
- Benchmark:
- Onboarded and standardized evaluation across multimodal benchmarks: Maverix, VMME, MMMU, MLVU, LongVideo Reasoning, HourVideo, VCR-Bench, VideoHolmes, VideoMathQA.
- Built knowledge-enriched evaluation suites (logo, BioCLIP concepts, landmarks) to measure fine-grained concept alignment.
- AI Research Scientist Intern, June 2024 – Nov. 2024
- GenAI, Meta, Menlo Park
- Mentor: Dr. Zecheng He
- Proposed and built a non-Markov conversational image generation framework that conditions on full interleaved text–image history, enabling rollback editing and name-based identity retrieval across multi-turn dialogues.
- Designed two non-Markov multi-round datasets: rollback-style non-markov multi-round editing and name-based non-markov multi-round personalization)
- Improved identity consistency via token-level caching, a DiT-based high-fidelity image detokenizer, and multi-stage instruction fine-tuning, significantly enhancing multi-round editing reliability and personalization stability.
- One paper accepted by WACV 2026 [10]
- Advanced Analytic Intern, May 2023 – Aug. 2023
- OXY Petroleum, Houston
- PI: Dr. Nikolaos Mitsakos, Dr. Ahmad Mustafa, Dr. Daniel De Lilla
- Adopted segmentation network for seismic horizon tracking task, seeking to alleviate the interpretation burden on geologists dealing with 3D volumetric data.
- Improved the performance with label trick, multi-inline input and ConvLSTM;
- Predicted 80% of horizon within 3-pixel-tolerance using only 1% manual label.
- Computer Vision Research Intern, Nov. 2020 – Aug. 2021
- AI Lab, ByteDance, Beijing
- PI: Dr. Zehuan Yuan
- Trial-and-error: Proposed and implemented many ideas in computer vision field, especially low-level vision.
- For example, using knowledge distillation as implicit prior for super-resolution;
- Searching the best simulation degradation for real world super-resolution.
- One paper published in ICCV 2021 [4].
Research Experience
- Auto-aggressive conversational image generation, June 2024 – Nov. 2024
- Meta GenAI @ Menlo Park, Mentor: Dr. Zecheng He
- Introduced conversational image generation, a new paradigm that enables multi-round personalized image generation through natural language interaction.
- Identified a key bottleneck in SEED-X MLLM framework—the visual detokenizer’s inability to preserve fine-grained facial identity. Thus, we replace its detokenizer with a personalization-enhanced DiT and propose a multi-stage instruction fine-tuning strategy that balances identity preservation and editability.
- To support non-Markov multi-round personalization, we built a chat-history caching mechanism and constructed the first name-based multi-round personalization dataset from video clips, enabling the model to reason jointly over past text and image context.
- Experiments demonstrate state-of-the-art personalization performance among MLLM-based methods and show, for the first time, that MLLMs can generate personalized images across multiple conversational turns.
- This paper has been accepted by WACV 2026 [10].
- Diffusion based fundus image synthesis for imbalanced diabetic retinopathy grading, Mar. 2024 – Apr. 2025
- Video Processing Lab @ UCSD, PI: Dr. Truong Q. Nguyen
- Proposed a diffusion-based data synthesis framework for imbalanced diabetic retinopathy grading by fine-tuning a text-to-image diffusion model on fundus images using a DreamBooth-style strategy with semantic-aware supervision.
- Introduced a semantic quality evaluation and self-supervised explicit class conditioning scheme to generate training samples that are diagnostically useful rather than merely visually realistic.
- Experiments show a substantial improvement in balanced accuracy (≈66.8% → 74.2%), outperforming oversampling and naive diffusion while reducing class-prior bias.
- This work has been published in MICCAI 2025 [9].
- 2D and 3D Retina OCT Angiography images classification, Dec. 2021 – Feb. 2024
- Video Processing Lab @ UCSD, PI: Dr. Truong Q. Nguyen
- Detect and differentiate active and inactive choroidal neovascularization (CNV) in different stages of age-related macular degeneration (AMD) using Optical Coherence Tomography Angiography (OCTA) scans.
- The challenges are small dataset with unbalanced distribution and potential retina layer segmentation errors.
- For Heidberg OCTA instrument, we conducted the first study to exclusively use OCTA data for AMD stage grading, demonstrating the modality’s diagnostic potential: Identified segmentation errors in retinal layers as a critical challenge for accurate classification using 2D OCTA projections; Proposed a novel approach of analyzing 3D OCTA volumes with 2D convolutional neural networks trained with additional projection supervision; Showed superior AI performance compared to human experts in grading AMD stages. One paper has been published in ICCVW 2023 [6]. Please also refer to our clinical paper [5].
- For cross-instrument data, Heidberg and Optovue, we explored cross-instrument disease classification by training DNN models on separate and combined datasets from both instruments: Employed style transfer techniques to generate cross-domain samples; Introduced a novel class-conditioned CycleGAN that integrates class-related constraints during training to optimize the generated samples for downstream classification tasks. One paper has been published in EMBC 2024 [7]. Please also refer to our clinical paper [8].
- Domain Adaptation based Unpaired Super-Resolution, Nov. 2020 – Apr. 2021
- ByteDance AI Lab @ Beijing, Mentor: Dr. Zehuan Yuan
- Formulated unpaired SR training as a feature-level domain adaptation problem, where the given LR images can be regard as the inputs in target domain and the provided HR images can be seen as label in source domain.
- Adopted adversarial-based feature distribution alignment to close the gap between source and target feature domains, and proposed several feature domain regularizations to achieve better aligning performance as well as preserve image details for the downstream SR task.
- This work has been published in ICCV 2021 [4].
- Tradeoff in Signal Restoration, Mar. 2019 – July. 2020
- MOE-Microsoft Key Laboratory of Multimedia Computing and Communication @ USTC, PI: Dr. Dong Liu
- Analyzed the relationship among signal fidelity, perceptual naturalness and semantic quality in the image restoration tasks. Demonstrated a tradeoff among the three metrics theoretically and experimentally.
- This work has been published in NeurIPS 2019 [3].
- Video Super-Resolution and Inverse Tone-Mapping, Nov. 2019 – Jan. 2020
- MOE-Microsoft Key Laboratory of Multimedia Computing and Communication @ USTC, PI: Dr. Dong Liu
- Participated in the 4K-HDR track of the first National Artificial Intelligence Challenge (NAIC).
- After passing the preliminary (super-resolution on noisy video) and the semi-final (video super-resolution + enhancement) round, we entered the final round with the best performance.
- During the final round, we completed the video super-resolution + SDR to HDR task using the specified machine within the specified time. Through subjective and objective evaluation and on-site defense, we finally won the runner-up prize with ¥500,000 bonus.
- Video Super-Resolution Tailored for Action Recognition, Sept. 2017 – Mar. 2019
- MOE-Microsoft Key Laboratory of Multimedia Computing and Communication @ USTC, PI: Dr. Dong Liu
- Investigated the video super-resolution (SR) problem for facilitating video analytics tasks rather than for visual quality.
- Tailored for two-stream action recognition networks, we developed SR methods for the spatial and temporal recognition respectively. On the one hand, we proposed an optical-flow guided weighted MSE to emphasize the reconstruction of moving objects. On the other hand, we proposed a siamese network training strategy in order to guarantee the temporal continuity between consecutive frames.
- This work has been published in ICCV 2019 [2].
- Text Image Super-Resolution Tailored for OCR, Sept. 2016 – June 2017
- MOE-Microsoft Key Laboratory of Multimedia Computing and Communication @ USTC, PI: Dr. Dong Liu
- Developed text image SR method to help optical character recognition (OCR).
- We proposed an edge-based loss function for SR training and conducted model combination to further improve the performance. Besides, we also developed an image padding method to refine the image boundaries during SR.
- This work has been published in VCIP 2017 [1].
Publications
Also see pulication page.
[1] Haochen Zhang, Dong Liu $^*$, Zhiwei Xiong. CNN-based Text Image Super-Resolution Tailored for OCR, In VCIP, St. Petersburg, FL, USA. Dec.10-13, 2017.
[2] Haochen Zhang, Dong Liu $^*$, Zhiwei Xiong. Two-Stream Action Recognition-Oriented Video Super-Resolution, In ICCV, Seoul, South Korea. Oct.27-Nov.2, 2019.
[3] Dong Liu $^*$, Haochen Zhang, Zhiwei Xiong. On the Classification-Distortion-Perception Tradeoff, In NeurIPS, Vancouver, Canada. Dec.8-14, 2019.
[4] Wei Wang $^\dagger$, Haochen Zhang $^\dagger$, Zehuan Yuan $^*$, Changhu Wang. Unsupervised Real-World Super-Resolution: A Domain Adaptation Perspective, In ICCV, Virtual. Oct.11-17, 2021.
[5] Anna Heinke, Haochen Zhang, Daniel Deussen, Carlo Galang, Alexandra Warter, Fritz Kalaw, Dirk-Uwe Bartsch, Lingyun Cheng, Cheolhong An $^* $, Truong Nguyen $^* $, William Freeman. Artificial intelligence for OCTA-based disease activity prediction in age-related macular degeneration, In RETINA, 2022.
[6] Haochen Zhang, Anna Heinke, Carlo Galang, Daniel Deussen, Bo Wen, Dirk-Uwe Bartsch, William Freeman, Truong Nguyen $^* $, Cheolhong An $^* $. Robust AMD Stage Grading with Exclusively OCTA Modality Leveraging 3D Volume, In ICCVW, Paris, France. Oct.2-6, 2023.
[7] Haochen Zhang, Anna Heinke, Krzysztof Broniarek, Carlo Galang, Daniel Deussen, Ines Nagel, Katarzyna Michalska-Małecka, Dirk-Uwe Bartsch, William Freeman, Truong Nguyen $^* $, Cheolhong An $^* $. OCTA-based AMD Stage Grading Enhancement via Class-Conditioned Style Transfer, In EMBC, Orlando, FL, USA. July 15-19, 2024.
[8] Anna Heinke, Haochen Zhang, Krzysztof Broniarek, Katarzyna Michalska-Małecka, Wyatt Elsner, Carlo Galang, Daniel Deussen, Alexandra Warter, Fritz Kalaw, Ines Nagel, Akshay Agnihotri, Nehal Mehta, Julian Elias Klaas, Valerie Schmelter, Igor Kozak, Sally L Baxter, Dirk-Uwe Bartsch, Lingyun Cheng, Cheolhong An $^* $, Truong Nguyen $^* $, William Freeman. Cross-instrument optical coherence tomography-angiography (OCTA)-based prediction of age-related macular degeneration (AMD) disease activity using artificial intelligence, In Scientific Reports, 2024.
[9] Haochen Zhang, Anna Heinke, Ines Nagel, Dirk-Uwe Bartsch, William Freeman, Truong Nguyen $^* $, Cheolhong An $^* $. Class-Conditioned Image Synthesis with Diffusion for Imbalanced Diabetic Retinopathy Grading, In MICCAI, Daejeon, South Korea, Sept.23-27, 2025.
[10] Haochen Zhang, Animesh Sinha, Felix Juefei-Xu, Haoyu Ma, Kunpeng Li, Zhipeng Fan, Xiaoliang Dai, Tingbo Hou, Peizhao Zhang, Zecheng He $^*$. Conversational Image Generation: Towards Multi-Round Personalized Generation with Multi-Modal Language Models, In WACV, Tucson, AZ, USA, March 6-10, 2026.
$^*$ denotes my advisor, $^\dagger$ denotes equal contribution co-author
