DATA ANALYTICS · MACHINE LEARNING · DEEP LEARNING · BUSINESS & PRODUCT

Yi Pan

I hold a Master of Data Science from Rice University. I have worked on computer vision, ML services, and customer analytics in US teams, as well as product support and supply chains for North America. I start with business needs, use data and technology to investigate problems, and turn tested solutions into product and business recommendations.

Explore my projects

ACADEMIC FOUNDATION

Education

  • Rice UniversityU.S. Top 20
    Master of Data Science
  • Drexel UniversityU.S. Top 100
    Bachelor of Science · Data Science
  • Lanzhou UniversityProject 985
    Bachelor of Engineering · Computer Science & Technology

U.S. labels: U.S. News 2027 · National Universities

Projects & Experience

01

AUDIENCE & BUSINESS

HGOAudience-to-Donor Strategy

University–industry collaboration

From Audience to Potential Donors.

Working with Houston Grand Opera, we combined attendance, donation, and survey records to study first and repeat giving and inform who to engage and when.

Donor conversionSurvival analysisCustomer analytics
Houston Grand Opera donor conversion project poster presented by the Rice teamView presentation poster
≈13,000

Customer Records

Before preprocessing

4

Data Sources

Survey · Ticketing · Donations · Marketing

46

Features

Background · Attendance · Engagement

After preprocessing, we built task-specific samples for first and repeat donations, combining customer background, attendance, and time-to-event records.

Audience Geography · Analysis & Visualization

Mapping ZIP codes turns many separate customer records into a view of geographic reach. The audience extends across the US, so a Houston-only profile would miss part of it. This view informs regional grouping; attendance and donation records are then needed to evaluate geographic differences.

Geographic distribution of customer ZIP codes · Warmer areas indicate more concentrated records · © OpenStreetMap contributors
Geographic distribution of customer ZIP codes · Warmer areas indicate more concentrated records · © OpenStreetMap contributors

Why Model When Someone Gives?

Outreach takes time. A new attendee and a long-standing audience member may need different engagement schedules. Survival analysis keeps the waiting time and can use records where no donation has been observed yet. Cox offers interpretable feature associations; Random Survival Forest tests more complex combinations of features.

Task and ModelTraining C-indexEvaluation C-index
First donation · Cox, full feature set0.7310.717
First donation · Random Survival Forest0.8000.733
Second donation · Cox, before feature selection0.6030.581

The study used an 80/20 training–evaluation split. These are exploratory results: the evaluation set also informed feature and parameter choices, so it was not an untouched final validation set.

Reading the Results

C-index: Who Donates First?

If A donates before B and the model puts A first, that pair is correctly ordered. A score of 0.733 is roughly 73 correct out of 100 comparable pairs; 0.717 is about 72, and 0.581 about 58. A score of 0.5 is near random ordering—not a customer’s donation probability.

Cox: Which Features Relate to Earlier Giving?

The full first-donation model gives annual attendance a hazard ratio (HR) of 1.11. For 4 versus 3 annual visits, holding other features equal and before either person donates, the instantaneous donation rate is about 11% higher—not an 11-percentage-point probability increase.

Kaplan–Meier: Who Is Still Waiting?

It estimates a waiting curve from observed records. If 10 people are followed for a full year and 3 donate, the curve ends at 70%. Someone observed for only six months still contributes those six months; they are not treated as someone who will never donate.

About C-index and survival curves
First donation · Cox predictions · One line per customer. The horizontal axis is years of waiting; the vertical axis is the predicted chance of no first donation yet. Faster decline means an earlier predicted donation.
First donation · Cox predictions · One line per customer. The horizontal axis is years of waiting; the vertical axis is the predicted chance of no first donation yet. Faster decline means an earlier predicted donation.
Second donation · Cox predictions · Time starts at the first gift. A higher curve means a higher predicted chance of still waiting for the next gift.
Second donation · Cox predictions · Time starts at the first gift. A higher curve means a higher predicted chance of still waiting for the next gift.

Think of the curve as “how many are still waiting to give?” A height of 70% at year 2 is like expecting about 70 out of 100 similar customers still to be waiting, with about 30 having given. A faster drop means an earlier expected donation.

Audience Profile from Analysis and Model Evaluation

A Profile Summarized by the Project
  • Age55+
  • EducationBachelor’s or higher
  • Household Income$125k+ / year
  • LocationHouston
  • RelationshipSubscriber
  • EngagementHigh satisfaction & recommendation intent

A profile drawn from audience analysis and model findings to understand differences and discuss outreach.

What Do These Features Help Us Ask?
  • Depth of engagement

    Read attendance alongside subscription status to distinguish occasional visits from sustained involvement.

  • Quality of the relationship

    Use satisfaction and recommendation intent as separate features: enjoying an experience and recommending it are different signals.

  • Timing of engagement

    Study the wait for first and repeat donations to discuss when to build a relationship and when to follow up.

Engagement features combine attendance records with survey responses: annual visits, satisfaction, and recommendation intent are linked to each customer and analyzed alongside donation timing. The NPS-style recommendation response is an individual input, adding relationship quality to the frequency of attendance.

EXPLORE THE OUTPUT

Interactive Demo

Set a prospective donor’s profile, then compare one change at a time. The examples explain how customer relationships and the time window affect the simulated output.

University–industry collaboration: source data and models are private. This demo visualizes the project’s output format using illustrative rules and example values.

01 / Customer Profile
02 / Relationship with HGO

Fields reflect the project’s research and audience profile. Age increases in small steps in this demo only; the Cox results did not establish a positive age effect. Region bands are illustrative: ≤50, >50–150, and >150 km.

SIMULATED SCENARIO
Chance of a First Donation Within This Period
—

How the Time Window Changes the View

More time creates more opportunity for a first donation. The example keeps this profile unchanged.

Compare One Change

Why Does the Number of Years Matter?

Five illustrative propensity levels: very low <15%, low 15–<30%, medium 30–<50%, high 50–<70%, very high ≥70%. Comparisons explain the simulation’s rules, not measured causal effects.

Private university–industry collaboration repository. Access requires authorization.

↑ Project index
02

VISION & PHYSICAL SYSTEMS

CPIOverlapping Product Recognition

Software Engineering Internship · Sep 2022–Mar 2023

Recognize Overlap. Test the Retrofit.

A camera on the existing arm moves to the center of a product column. Detection models identify overlapping products; linear regression helps select the center-column detections before counting. Local tests establish the retrofit’s useful range under occlusion and dim light.

YOLOv7 · PyTorchYOLOv4 · TensorFlowOpenCVDockerAWS S3 / EC2Linear Regression
Device demonstration · 8 sec · silent
RETROFIT OBJECTIVE

Test how much useful vision can be added to an existing machine: position the camera, detect overlapping products, then validate counting depth on the physical prototype. The videos show the resulting local recognition workflow.

Mechanical Engineering

3D design · Arm retrofit · Camera installation

My Work

Camera control · Data pipeline · Models · Device integration

Why YOLO for This Prototype?

R-CNN was also tried. For a continuous camera feed, the selection focused on recognition under overlap, training iteration, and local deployment. YOLO’s single-stage detection pipeline suited this workflow; YOLOv7 with PyTorch was retained after comparison with YOLOv4.

Model Comparison

A Different Model. A Stronger Result.

Same dataset · Same EC2 · Equal epoch count

YOLOv7 · PyTorch≈99%

Final model · Test accuracy

>2×YOLOv7 / PyTorch vs. YOLOv4 training speed

Approximate historical test accuracy. Training speed compares the same data, image size, epoch count, and EC2 configuration. Model and framework changed together.

Implementation Workflow

  1. Device & Local Scripts

    Data Collection
    1. 01
      Position & Capture

      Move the camera to the target column’s center; capture images with Python / OpenCV.

    2. 02
      Prepare Images

      Organize images and labels for object-detection experiments.

  2. AWS Cloud Platform

    Cloud Training
    1. 03
      Store & Connect

      Amazon S3 stores the images; an Amazon EC2 GPU instance runs training on Linux. SSH keys provide secure remote access.

    2. 04
      Train & Select

      YOLOv4 with TensorFlow for earlier experiments; YOLOv7 with PyTorch for the final model.

  3. Physical Validation

    Local Deployment & Limits
    1. 05
      Integrate the Model

      Move the selected model from Ubuntu to Windows and connect the camera and positioning system.

    2. 06
      Filter & Count

      Use linear regression to help select center-column detections, then count the retained products and inspect the result on the prototype.

Cloud compute supported training; inference was integrated into the local prototype. Validation covers both recognition and the physical conditions needed for it.

DETECTION → COLUMN FILTER → COUNT

Count the Target Column

A camera frame can include products in neighboring columns. After detection, linear regression supports the column-selection step, filtering the results to the center column before counting. This keeps nearby products from being included in the target count.

YOLO · Find productsLinear Regression · Select the column
Center column · CountedAdjacent columns · ExcludedSelection concept · Not an experimental output

The Hard Part: Seeing Behind the Front Item

Products share the same line of sight. The task is to separate partially visible items, rather than count eight clearly separated objects.

PARTIAL OCCLUSION · SCHEMATICEdges remain visible on alternating sides
01020304050607–08 · HiddenREAR / LESS LIGHTCamera at the column center
Small left–right offsets expose different edges and labels. In the tested setup, the first six items were recognizable; the rear two had almost no usable visible area.
What the Model Must Learn

Use the remaining visible edges and appearance to distinguish overlapping products. Small exposed areas and dim light make the rear items harder.

Where the Limit Appears

Almost no visible pixels remain for the last two items. This identifies where lighting or viewpoint changes may be needed.

Prototype demonstration · 13 sec · silent

Company-owned work. Demonstration videos show the prototype output; source code, internal datasets, and non-public implementation details are not shared.

↑ Project index
03

GRAPHS & PREDICTION

NYCTraffic Prediction

Team course project · 2022

Connected Roads. Better Context.

Combining Uber Movement hourly speeds with OpenStreetMap, the model uses each road’s history and connected roads to predict the next hour’s speed.

Graph attentionPyTorchTime seriesExploratory analysis
New York City road network used in the traffic projectNew York City road network · Project visualization
Graph nodes
21,947
Graph edges
74,783
Temporal Inputs per Node
24

Why Look Beyond One Road’s History?

A road’s past speed describes its own rhythm; its connections show where neighboring traffic may matter. Graph theory represents those connections, and GAT learns how much neighbor information to use. Each node receives 24 hourly speed values: one measured variable across 24 time steps. The study also compared Prophet and Random Forest, testing the value of a network-aware approach.

Betweenness centrality distributions across congestion groups

Network Structure Offers Context

This comparison groups roads by congestion and examines how often each lies on shortest routes. The overlap shows that structure alone cannot explain traffic conditions; hourly speed histories are needed too.

MODEL OUTPUT

Predicted vs. Observed Speed

Red shows predictions and blue shows observations. Their broad rises, falls, and value ranges are similar, with several peaks and troughs occurring around the same time. The model captures much of this node’s overall pattern.

The prediction is smoother, and some sharp peaks and sudden drops remain imperfect. Look at both whether it follows the trend and where it misses abrupt changes.

One road-network node during the test period · Red: predicted · Blue: observed
One road-network node during the test period · Red: predicted · Blue: observed
INTERACTIVE SIMULATION

What Changes a Road’s Forecast?

Choose a scenario and a road. Edit an hourly speed or a connected road, then see how the next-hour estimate changes.

Start with an example
SelectedDirect neighbor

24-hour mean
Direct connections
Road D

Each hour is one input value. Changing an earlier hour changes the history without necessarily changing the latest speed.

60300−23 hNow+1 h
Blue: entered historyGold: next-hour estimate
Edit a Neighbor’s Latest Speed

Only directly connected roads contribute to this local information exchange.

mph

What Did Neighbor Information Change?
Own history only
With neighbors

Which Neighbors Get More Weight?

The project uses 24 hourly speeds per node and road connections to predict the next hour. This simulation makes those inputs editable with example data and illustrative calculations. The two estimates above explain the demo’s information flow, not a model-comparison score. Congestion cutoffs follow the study: 15.91 / 20.88 / 33.80 mph.

MIT License · Team: Harry Zhao, Yi Pan, Yantian Ding. Static figures and experiment metrics come from the team report; this was a course project, not a live traffic service.

↑ Project index
04

DATA & APPLICATIONS

StarTreeConfiguration Recommendations

Software Engineering Internship · May–Aug 2024

From Operating Load to Resource Advice.

Combine historical operating metrics and web data to recommend infrastructure settings such as CPU core count, giving capacity planning a data-informed starting point.

Pinot SQL · WgetScikit-learn · FlaskREST · DockerPrompt Engineering
01

Data Preparation

SQL

Historical operating metrics

Kafka → Apache Pinot
Wget

Web data

Retrieve reference inputs

Clean, align and combine inputs

Grafana · SQL queries & historical load visualization

02

Prediction Service

Scikit-learn

Random Forest Regression

FlaskREST API

Make predictions available to the application

03

Application

Java · Spring Boot

Call the prediction API with request inputs

RECOMMENDATIONCPU Cores

Initial infrastructure configuration

DELIVERYDocker·GitHub ActionsContainerization & delivery workflow

INFRASTRUCTURE CONTEXT

What Is Running—and What Needs Capacity?

Pinot separates query routing, data processing, and cluster management. The examples below explain the roles and signals behind capacity planning.

01
Broker

Routes SQL queries and merges results.

Query volume · Response time
02
Server

Stores data segments and runs queries.

CPU · Memory · Ingestion delay
03
Controller

Manages metadata and cluster resources.

Segment health · Cluster state
CPU Usage ≠ CPU Cores

Usage describes how busy the allocated CPU is; core count describes how much computing capacity is assigned. Historical load helps inform the starting configuration, rather than treating every workload the same.

Component and metric examples: Apache Pinot · StarTree Capacity Planning

The model uses historical examples to recommend initial settings such as CPU core count. An API brings that advice into the application, helping connect observed load with resource and cost choices.

My Contribution

Built data preparation, a Flask prediction service, Docker packaging, and REST integration with the Java backend. Also explored prompts for a separate internal AI task.

SQL + Grafana

Queried historical metrics with SQL in Grafana and visualized CPU usage to understand operating load alongside the configuration-recommendation work.

INTERNAL AI EXPLORATION

Prompt Engineering

Configuration scenario · Illustration

Designed prompts and iterated on model responses for an internal AI task. This configuration example shows how business requirements can be translated into a clear input and output contract.

QPS
Queries per second
Throughput
Incoming data rate
Latency
Target response time

Specify the workload and available context before requesting a configuration.

JSONMissing-context example
{
  "cpu_cores": null,
  "needs_clarification": true,
  "missing_inputs": [
    "query_complexity",
    "data_volume"
  ]
}
JSON onlyRequired: cpu_coresClarify missing inputs

Require a positive integer for cpu_cores when enough information is available; otherwise return null and the missing inputs.

Company project. The architecture summarizes the documented workflow. Internal code and data are not public; the separate prompt-engineering exploration did not reach deployment.

↑ Project index

Technical Practice+Product Understanding

Hands-on work in data analysis, model training, and system integration helps me connect business needs with what technology can deliver: identify the problem, compare solutions, and test which capabilities are useful in a product.

At SUMEC / FIRMAN, my work for North America also involved product-development coordination, cost considerations, supply-chain delivery, and after-sales feedback. I follow generative AI and multimodal models, with an interest in how new capabilities can meet user needs and improve products.

Understand the Business

Python · pandas · SQL

Use customer behavior, business processes, and data to identify problems, define useful measures, and support recommendations.

Test the Technology

PyTorch · TensorFlow · Scikit-learn · OpenCV

Compare models and implementations through experiments, drawing on computer vision, graph models, and Transformers.

Inform the Product

Flask · Git · Docker · AWS S3 / EC2

Connect models to real workflows and test their limits to inform features, retrofit options, and resource choices.

Contact

  • AI Product Manager
  • Machine Learning Engineering
  • Business & Supply-chain Analytics
  • Data Analytics

Interested in work that connects technology, products, and business needs. Highly willing to travel internationally and take part in overseas projects; comfortable working in English and Chinese.