Iβm a Data Science student at UC San Diego, interested in machine learning systems, distributed computing, and high-performance computing.
My work spans LLM and retrieval systems, distributed inference, GPU-accelerated software, and data/ML infrastructure across cloud and HPC environments.
- Building and deploying LLM-powered retrieval and backend systems
- Exploring GPU acceleration, distributed inference, and ML systems performance
- Developing systems projects in Rust, CUDA, and modern GPU APIs
- Benchmarking and investigating large-scale scientific simulation performance
- Toaster β Rust/wgpu GPU path tracer with BVH acceleration, animation, HDR lighting, physics, and deterministic benchmarking.
- Wasserstein Hypergraph Alignment β Hypergraph neural network for 3D point-cloud alignment using Wasserstein aggregation.
LLM Engineer Intern @ San Diego Supercomputer Center β San Diego, California
Mar 2026 β Present
- Build end-to-end LLM applications integrating data ingestion, retrieval, backend services, and user-facing systems.
- Develop information retrieval systems using Neo4j and Microsoft GraphRAG across heterogeneous datasets.
- Implement unit tests and CI/CD workflows for applications deployed in an on-premises Linux/HPC environment.
Data Science Intern @ Cloocus (Microsoft Azure Partner) β Kuala Lumpur, Malaysia
Jul 2025 β Sep 2025
- Standardized SQL datasets across 5+ sources, improving latency by 35% and reducing inconsistencies by 30%.
- Developed RAG systems over heterogeneous document sources for low-latency, citation-aware queries.
- Built real-time ingestion pipelines using SQL, Azure Blob, and REST APIs.
- Containerized and deployed ML services on Azure, improving deployment speed by 20%.
IEEE Supercomputing (UCSD) β Officer
- Led the D-LLaMA sub-team at SBCC 2025 across 16 Raspberry Pis; placed 2nd in D-LLaMA and 3rd overall.
- Captained UCSD and led D-LLaMA at SBCC 2026; placed 3rd in D-LLaMA and 3rd overall.
- Automated LLaMA rebuilds, parallel cluster deployment, systemd worker recovery, and Vulkan environment setup.
- Benchmarked D-LLaMA GPU acceleration and prioritized the faster CPU/distributed path.
- Ran and optimized MLPerf and scientific workloads across multi-GPU and HPC systems.
SBCC 2025 Results Β· SBCC 2026 Results
Rust, wgpu, WGSL, GPU Computing
- Built and integrated a Rust/wgpu GPU renderer supporting glTF meshes, HDR lighting, animation, rigid-body physics, and progressive rendering.
- Integrated GPU execution, scene evaluation, benchmarking, preview, and export infrastructure across the renderer.
- Added glTF collider proxies for physics-driven scenes and measured up to 89.3% faster rendering on larger RTX 2080 Ti workloads.
C++, CUDA, Linux, Docker
- Developed a modular GPU-accelerated image processing system with CUDA kernels for grayscale, box blur, and Sobel edge detection.
- Benchmarked CPU and GPU execution on RTX 2080 Ti hardware to analyze parallel performance.
- Structured the project into separate kernels and build components for easier testing and extension.
Python, PyTorch, PyTorch Geometric
- Developed a hypergraph neural network pipeline for 3D point-cloud alignment under rigid, affine, and noisy transformations.
- Replaced HyperGCT mean pooling with Wasserstein-based aggregation and evaluated it on FAUST and PartNet.
- Improved F1 from 0.782 to 0.956 on the hardest setting, reaching up to 0.993 overall.
Python, PyTorch, Graph Neural Networks
- Implemented and benchmarked MPNNs, GraphGPS, and graph Transformers under controlled settings.
- Evaluated models on synthetic graph algorithms and ZINC-12k molecular regression.
- Showed graph-native models outperform sequence Transformers under distribution shift.
Python, Pandas, scikit-learn, Matplotlib, Seaborn
- Analyzed professional League of Legends match data to study how gold distribution and role dynamics affect player performance.
- Built predictive models and evaluated performance across different game contexts.
- Performed fairness and error analysis to assess model consistency across player roles and scenarios.
Python, PyTorch, Raspberry Pi Zero
- Built and deployed an LSTM-based sentiment analysis model for real-time inference on a Raspberry Pi Zero.
- Tuned model hyperparameters using Bayesian optimization.
- Achieved 97% validation accuracy and earned 2nd place at an IEEE competition.
Python, Matplotlib
- Implemented GD, AdaGrad, RMSProp, AdaDelta, and Adam from scratch.
- Visualized optimizer trajectories over complex loss surfaces using Matplotlib animations.
- Compared convergence behavior across different optimization algorithms.
Python, scikit-learn
- Developed a deepfake detection pipeline using facial landmark features.
- Trained and evaluated classification models using cross-validation and quantitative metrics.
- Visualized model behavior and prediction results to analyze detection performance.
Python, TensorFlow, CUDA
- Built an image classification pipeline under constrained compute resources.
- Investigated preprocessing and training strategies to improve model efficiency and accuracy.
- Used CUDA acceleration to achieve 95% validation accuracy.
Python, LSTM
- Built a 2D physics simulation engine to generate motion and trajectory data.
- Trained an LSTM model to predict object trajectories from simulated sequences.
- Earned Bronze at KSEF for the project.
π§ cypark1516@gmail.com
πΌ LinkedIn
π» GitHub


