DATA ANALYTICS · MACHINE LEARNING · DEEP LEARNING

Yi Pan

Data-mindedHands-onBusiness-curious

I like understanding how things work—and how they can work better for a business. My background spans data science, hands-on engineering in US teams, and supply-chain coordination for North America. I’m interested in AI applications and data products that turn a real need into something useful.

Explore my projects

ACADEMIC FOUNDATION

Education

  • Rice UniversityMaster of Data Science
  • Drexel UniversityBachelor of Science · Data Science
  • Lanzhou UniversityBachelor of Engineering · Computer Science & Technology
Open to AI Application Product, Data Product & Business Analytics roles

Technical Depth+Business Context

A data science foundation, experience building working systems, and an interest in the decisions those systems support.

At SUMEC / FIRMAN, four months of work on North American product support and supply-chain coordination gave me practical exposure to customer feedback, parts orders, and delivery constraints.

Analyze

Python · pandas · SQL

Customer analysis, metric definitions, and evidence-based recommendations.

Model and Architecture

PyTorch · TensorFlow · Scikit-learn · OpenCV

Machine learning, Transformers, graph theory, and computer vision.

Integrate

Flask · Git · Docker · AWS S3 / EC2

Connecting models to applications and collaborating across engineering disciplines.

Projects & Experience

Customer decisions. Physical systems. Connected services. Four projects, viewed through their results and the work behind them.

01

AUDIENCE & BUSINESS

Houston Grand Opera

University–industry collaboration

Understanding when opera audiences become donors.

A Rice D2K collaboration with Houston Grand Opera: connecting audience profiles, donation timing, and fundraising decisions through survival analysis.

Donor conversionSurvival analysisCustomer analytics
Houston Grand Opera donor conversion project poster presented by the Rice teamView presentation poster
Customers before cleaning & filtering
~13,000
First-donation RSF · Evaluation C-index
0.733

The source data covered approximately 13,000 customers across survey, ticketing, donation, and marketing records. Unusable or unsuitable records were excluded during preparation; the experiments below use their eligible subsets, not all 13,000 customers. The analysis used 46 features, including age and income.

THE DECISION

Who should the team understand and engage next?

I worked across data integration, EDA, survival modeling, evaluation, and recommendations in the Rice–HGO team, and presented the findings through a poster and live Q&A.

4

Data Sources

Survey · Ticketing · Donations · Marketing

46

Features

Age · Income · Attendance · Subscription

3

Analysis Methods

Kaplan–Meier · Cox · Random Survival Forest

From Customer Profile to Outreach Plan

A Profile to Explore
  • Age55+
  • EducationBachelor’s or higher
  • Household Income$125k+ / year
  • LocationHouston
  • RelationshipSubscriber
  • EngagementHigh satisfaction & recommendation intent

A group-level profile to guide outreach exploration, not a description of every donor or an individual eligibility rule.

Recommended Actions
  • Frequent Attendance

    Give repeat visitors relevant follow-up and consider annual visit frequency when planning outreach.

  • Subscription Relationships

    Build on existing subscriber relationships and strengthen subscriber benefits.

  • Strong Recommendation Intent

    Support engaged customers with a good experience and opportunities to recommend HGO to others.

NPS (Net Promoter Score) summarizes how willing customers are to recommend an organization. Here, recommendation responses help describe engagement; they are distinct from overall satisfaction.

These are proposed actions informed by the analysis; resulting fundraising gains were not measured in this project.

Where Are the Customers?

Mapped customer ZIP codes to explore their geographic distribution across the United States. The heatmap helps describe audience reach; warmer areas indicate stronger concentrations in the displayed data, not higher donation rates.

Customer distribution by ZIP code · EDA output · Basemap © OpenStreetMap contributors
Customer distribution by ZIP code · EDA output · Basemap © OpenStreetMap contributors

When Might a Customer Donate?

Survival analysis studies how long it takes for an event to happen. Here, the event is a first or second donation. Cox regression gives an interpretable statistical model; Random Survival Forest (RSF) combines decision trees to capture more complex patterns.

Task and ModelTraining C-indexEvaluation C-index
First donation · Cox, full feature set0.7310.717
First donation · Random Survival Forest0.8000.733
Second donation · Cox, before feature selection0.6030.581

Saved notebook outputs, rounded to three decimals, from 80/20 splits. These are exploratory results: evaluation data were also consulted during feature and parameter selection. They are not a fresh, untouched final test. The poster’s second-donation score of 0.61 is from a full-data fit; feature-removal experiments reached about 0.625 on the reused evaluation split.

Reading the Results

C-index · Which Customer Donates Earlier?

Measures how well the model orders donation times among customer pairs that can be compared. 0.5 is chance-level ordering; 1.0 is perfect ordering. A score of 0.733 is roughly 73 out of 100 comparable pairs ordered correctly, with ties receiving partial credit. It does not mean 73.3% of customers will donate.

Survival Probability · Not Yet Donated

In these plots, “survival” means the donation has not happened yet. A value of 0.7 at year 2 means an estimated 70% chance of no donation by then; 1 − 0.7 = 30% is the estimated chance of a donation by that time. A faster fall indicates earlier predicted giving.

About C-index and survival curves
First donation · Predicted curves for five selected customers from the reduced-feature Cox model.
First donation · Predicted curves for five selected customers from the reduced-feature Cox model.
Second donation · Model prediction for an example customer profile.
Second donation · Model prediction for an example customer profile.

These curves illustrate model predictions, rather than observed conversion rates.

Interactive Demo · Rule-Based Simulation

From Customer Features to Decision Support

Explore a customer scenario, then filter a 500-customer simulated audience. This demo illustrates how customer analysis can become an input-and-output workflow and a BI view for fundraising discussions.

The original university–industry project uses private data and models that are not served here. This portfolio demo uses 500 synthetic customers and illustrative rules informed by the project discussion: little attendance and greater distance reduce simulated donation propensity. Rule strengths are set for demonstration, not fitted coefficients or outputs from the trained model.

Try a Customer Scenario

This scenario is a separate rule-based example, not a filter on the board below. Distance is an illustrative input, not a verified distance coefficient from the trained model.

Explore the 500-Customer Simulation

Donation Participation by Attendance
What to Investigate Next

Metric Definitions
Audience
Customers in the selected region and subscription group. This synthetic cohort is fixed at the start of the demo year.
Donation Participation
Customers with at least one donation in the selected window ÷ customers in the selected group. This is not an outreach conversion rate.
Repeat-Donation Share
Customers with two or more donations in the window ÷ customers with at least one donation in that same window. This is not next-period retention.
Donation Amount
Total value of donations within the selected window, in simulated US dollars. Attendance bands also use that window.
↑ Project index
02

VISION & PHYSICAL SYSTEMS

Crane Payment Innovations

Software Engineering Internship · Sep 2022–Mar 2023

Testing a vision retrofit for an existing vending machine.

A camera-based retrofit study: move the camera to the target column, recognize products, and establish the practical limits of depth-wise counting before the company considers further investment.

YOLOv7 · PyTorchYOLOv4 · TensorFlow 2.0AWS S3 / EC2Vision Retrofit
Device demonstration · 8 sec · silent

Model Comparison

A Different Model. A Stronger Result.

Same dataset · Same EC2 · Equal epoch count

YOLOv7 · PyTorch98–99%

Final model · Test accuracy

>2×Training speed · Equal epochs

Historical test-accuracy estimates recalled from the project, not F1 or mAP. Training time compares the same epoch count, dataset, image size, and EC2 configuration. Model version and framework both changed; the gain cannot be attributed to architecture alone.

THE BUSINESS CASE

Can an existing machine gain useful vision capabilities through a retrofit? My role was to establish technical feasibility and operating limits; the company evaluated cost and the next product decision.

Mechanical Engineering

3D design · Arm retrofit · Camera installation

My Work

Camera control · Data pipeline · Models · Device integration

Implementation Workflow

  1. Device & Local Scripts

    Data Collection
    1. 01
      Position & Capture

      Move the camera to the target column’s center; capture images with Python / OpenCV.

    2. 02
      Prepare Images

      Organize images and labels for object-detection experiments.

  2. AWS Cloud Platform

    Cloud Training
    1. 03
      Store & Connect

      Amazon S3 for data; SSH key authentication to a Linux-based Amazon EC2 GPU instance.

    2. 04
      Train & Select

      YOLOv4 with TensorFlow 2.0 for earlier experiments; YOLOv7 with PyTorch for the final model.

  3. Local Prototype

    Device Integration
    1. 05
      Integrate the Model

      Move the selected model from Ubuntu to Windows and connect the camera and positioning system.

    2. 06
      Run & Inspect

      Run product recognition and counting on the prototype; inspect detections in the camera feed.

Cloud compute supported training; inference was integrated into the local prototype. The diagram summarizes the confirmed work rather than a detailed infrastructure topology.

What the Retrofit Test Established

In the tested camera position and product arrangement, the system could correctly recognize the first six items in a column that physically held eight. The two rear items were almost fully hidden behind the front products, with less available light.

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06
  7. 07
  8. 08

Front of column → Rear · Positions 1–6: recognized in the test · Positions 7–8: strongly occluded

Schematic of the observed setup, not an accuracy benchmark. Six of eight positions is not a 75% accuracy score.

What Limited the View?

Light and occlusion constrain what the camera can observe. Better model performance does not remove the need for a usable view of the rear products.

What Could the Company Decide?

Weigh the demonstrated capability against retrofit cost. Lighting or camera-angle changes are possible follow-up tests, not completed improvements.

Official YOLOv7 implementation
Prototype demonstration · 13 sec · silent

Company-owned work. Demonstration videos show the prototype output; source code, internal datasets, and non-public implementation details are not shared.

↑ Project index
03

GRAPHS & PREDICTION

NYC Traffic Prediction

Team course project · 2022

Learning how connected roads influence traffic speed.

New York City traffic-speed prediction using Uber Movement data, OpenStreetMap road structure, and a multi-head graph attention model.

Graph attentionPyTorchTime seriesExploratory analysis
New York City road network used in the traffic projectNew York City road network · Project visualization
Graph nodes
21,947
Graph edges
74,783
GAT · MAE (report)
7.1

Why a Graph?

I contributed throughout the three-person project: data preparation, EDA, graph construction, multi-head GAT implementation, and evaluation. Road connections define which neighbors can exchange information; attention learns their relative weights.

EDUCATIONAL SIMULATION

One Road. A Network of Signals.

Select a node. Change its recent history or its neighbors. Follow the information into a next-hour speed estimate.

SelectedDirect neighborOther node

Node D

Controls shape a 24-hour example history. The neighbor adjustment applies to connected nodes only.

Illustrative Speed Estimate

—mph

Weighted Neighbor Speed
−23 hLatest hour

Each node carries 24 hourly speeds. Edges determine which nodes can share information.

h′i = Σj∈N(i) αij hjΣ α = 1

Compare history features → normalize scores → weight neighbor information. The demo uses fixed summary features; the report’s GAT learns its feature projections from data.

Demo for understanding the project’s outputs: example speeds and a simplified aggregation rule, rather than reported experiment results. Next-hour estimate = 65% latest speed + 25% weighted neighbor speed + 10% history mean + a trend adjustment. Congestion cutoffs follow the report (15.91 / 20.88 / 33.80); this demo uses mph, matching the source speed_mph_mean field.

Data Preparation & Network Structure

Preprocessing workflow · Report Figure 3
Preprocessing workflow · Report Figure 3

Build a graph from road identifiers, select a better-covered subgraph, then align hourly speeds and fill missing values. Filled speeds are assumptions, not additional observations.

Betweenness centrality by congestion group · Report Figure 24
Betweenness centrality by congestion group · Report Figure 24

Betweenness describes how often a road lies on shortest paths between other nodes. Comparing congestion groups explores structural context for a graph model. The distributions overlap; the plot alone does not establish causation or predictive validity.

MODEL OUTPUT

Predicted vs. Observed Speed

The report records a GAT MAE of 7.1 and RMSE of 9.55 (Table II), alongside a prediction-versus-observation plot for the final seven days. Values are quoted as reported; the table does not specify their units or scaling. The 744 time steps include filled gaps, rather than 744 measured observations for every node.

MAE is the average size of prediction errors. RMSE gives larger errors more weight. For the same target and scale, lower values are better; neither is an accuracy percentage.

Predicted and observed speed for one node in the report's test period.
Predicted and observed speed for one node in the report's test period.

Team: Harry Zhao, Yi Pan, Yantian Ding. Static figures and experiment metrics come from the team report; this was a course project, not a live traffic service.

↑ Project index
04

DATA & APPLICATIONS

StarTree

Software Engineering Internship · May–Aug 2024

From operational data to configuration recommendations.

A machine learning service for infrastructure capacity planning, connecting Python model inference with a Java application backend.

Pinot SQL · WgetScikit-learn · FlaskREST · DockerPrompt Engineering
DOCUMENTED INPUTS

Historical Metrics

Kafka → Apache Pinot

SQLGrafana · Query & Visualize

Web Data

Retrieve web content

Wget
DATA

Prepare & Integrate

Historical metrics + web inputs

MODEL SERVICE

Scikit-learn + Flask

Random forest regression · Initial configuration

APPLICATION

Java
Spring Boot

HTTP / REST

Configuration Recommendation

Docker · GitHub Actions

My Contribution

Built data preparation, a Flask prediction service, Docker packaging, and REST integration with the Java backend. Also explored prompts for a separate internal AI task.

SQL + Grafana

Queried historical metrics with SQL in Grafana and visualized CPU usage to understand operating load alongside the configuration-recommendation work.

Separate AI Exploration

Designed prompts and reviewed responses for an internal AI task; this exploration did not reach deployment.

Company project. The architecture summarizes the documented workflow. Internal code and data are not public; the separate prompt-engineering exploration did not reach deployment.

↑ Project index

NEXT CHAPTER

Contact

Open to AI Product Manager, Data Product, Data Analyst, and applied Data Scientist roles. Interested in work that connects technical delivery with business needs, including international teams and overseas projects.