DATA ANALYTICS · MACHINE LEARNING · DEEP LEARNING · BUSINESS & PRODUCT

Yi Pan

I have worked on computer vision, ML services, and customer analytics in US teams, as well as product support and supply chains for North America. I start with business needs, use data and technology to investigate problems, and turn tested solutions into product and business recommendations.

Explore my projects

Education

  • Rice UniversityU.S. News Top 20
    Master’s Degree · Data Science (MDS)
  • Drexel UniversityU.S. News Top 100
    Bachelor of Science · Data Science
  • Lanzhou UniversityProject 985
    Bachelor of Engineering · Computer Science & Technology

Projects & Experience

01

AUDIENCE & BUSINESS

HGOAudience-to-Donor Strategy

University–industry collaboration

From Audience to Potential Donors.

Working with Houston Grand Opera, we combined attendance, donation, and survey records to study first and repeat giving and inform who to engage and when.

Donor conversionSurvival analysisCustomer analytics
Houston Grand Opera donor conversion project poster presented by the Rice teamView presentation poster
≈13,000

Customer Records

Before preprocessing

4

Data Sources

Survey · Ticketing · Donations · Marketing

46

Features

Background · Attendance · Engagement

After preprocessing, we built task-specific samples for first and repeat donations, combining customer background, attendance, and time-to-event records.

Audience Geography · Analysis & Visualization

Mapping ZIP codes turns many separate customer records into a view of geographic reach. The audience extends across the US, so a Houston-only profile would miss part of it. This view informs regional grouping; attendance and donation records are then needed to evaluate geographic differences.

Geographic distribution of customer ZIP codes · Warmer areas indicate more concentrated records · © OpenStreetMap contributors
Geographic distribution of customer ZIP codes · Warmer areas indicate more concentrated records · © OpenStreetMap contributors

Why Model When Someone Gives?

Outreach takes time. A new attendee and a long-standing audience member may need different engagement schedules. Survival analysis keeps the waiting time and can use records where no donation has been observed yet. Cox offers interpretable feature associations; Random Survival Forest tests more complex combinations of features.

Task and ModelTraining C-indexEvaluation C-index
First donation · Cox, full feature set0.7310.717
First donation · Random Survival Forest0.8000.733
Second donation · Cox, before feature selection0.6030.581

The study used an 80/20 training–evaluation split. These are exploratory results: the evaluation set also informed feature and parameter choices, so it was not an untouched final validation set.

Reading the Results

C-index: Who Donates First?

If A donates before B and the model puts A first, that pair is correctly ordered. A score of 0.733 is roughly 73 correct out of 100 comparable pairs; 0.717 is about 72, and 0.581 about 58. A score of 0.5 is near random ordering—not a customer’s donation probability.

Cox: Which Features Relate to Earlier Giving?

The full first-donation model gives annual attendance a hazard ratio (HR) of 1.11. For 4 versus 3 annual visits, holding other features equal and before either person donates, the instantaneous donation rate is about 11% higher—not an 11-percentage-point probability increase.

Kaplan–Meier: Who Is Still Waiting?

It estimates a waiting curve from observed records. If 10 people are followed for a full year and 3 donate, the curve ends at 70%. Someone observed for only six months still contributes those six months; they are not treated as someone who will never donate.

About C-index and survival curves
First donation · Cox predictions · One line per customer. The horizontal axis is years of waiting; the vertical axis is the predicted chance of no first donation yet. Faster decline means an earlier predicted donation.
First donation · Cox predictions · One line per customer. The horizontal axis is years of waiting; the vertical axis is the predicted chance of no first donation yet. Faster decline means an earlier predicted donation.
Second donation · Cox predictions · Time starts at the first gift. A higher curve means a higher predicted chance of still waiting for the next gift.
Second donation · Cox predictions · Time starts at the first gift. A higher curve means a higher predicted chance of still waiting for the next gift.

Think of the curve as “how many are still waiting to give?” A height of 70% at year 2 is like expecting about 70 out of 100 similar customers still to be waiting, with about 30 having given. A faster drop means an earlier expected donation.

Audience Profile from Analysis and Model Evaluation

A Profile Summarized by the Project
  • Age55+
  • EducationBachelor’s or higher
  • Household Income$125k+ / year
  • LocationHouston
  • RelationshipSubscriber
  • EngagementHigh satisfaction & recommendation intent

A profile drawn from audience analysis and model findings to understand differences and discuss outreach.

What Do These Features Help Us Ask?
  • Depth of engagement

    Read attendance alongside subscription status to distinguish occasional visits from sustained involvement.

  • Quality of the relationship

    Use satisfaction and recommendation intent as separate features: enjoying an experience and recommending it are different signals.

  • Timing of engagement

    Study the wait for first and repeat donations to discuss when to build a relationship and when to follow up.

Engagement features combine attendance records with survey responses: annual visits, satisfaction, and recommendation intent are linked to each customer and analyzed alongside donation timing. The NPS-style recommendation response is an individual input, adding relationship quality to the frequency of attendance.

EXPLORE THE OUTPUT

Interactive Demo

Set a prospective donor’s profile, then compare one change at a time. The examples explain how customer relationships and the time window affect the simulated output.

University–industry collaboration: source data and models are private. This demo visualizes the project’s output format using illustrative rules and example values.

01 / Customer Profile
02 / Relationship with HGO

Fields reflect the project’s research and audience profile. Age increases in small steps in this demo only; the Cox results did not establish a positive age effect. Region bands are illustrative: ≤50, >50–150, and >150 km.

SIMULATED SCENARIO
Chance of a First Donation Within This Period
—

How the Time Window Changes the View

More time creates more opportunity for a first donation. The example keeps this profile unchanged.

Compare One Change

Why Does the Number of Years Matter?

Five illustrative propensity levels: very low <15%, low 15–<30%, medium 30–<50%, high 50–<70%, very high ≥70%. Comparisons explain the simulation’s rules, not measured causal effects.

Private university–industry collaboration repository. Access requires authorization.

↑ Project index
02

VISION & PHYSICAL SYSTEMS

CPIOverlapping Product Recognition

Software Engineering Internship · Sep 2022–Mar 2023

Recognize Overlap. Test the Retrofit.

A camera on the existing arm moves to the center of a product column. Detection models identify overlapping products; linear regression helps select the center-column detections before counting. Local tests establish the retrofit’s useful range under occlusion and dim light.

YOLOv7 · PyTorchYOLOv4 · TensorFlowOpenCVDockerAWS S3 / EC2Linear Regression
Device demonstration · 8 sec · silent
RETROFIT OBJECTIVE

Test how much useful vision can be added to an existing machine: position the camera, detect overlapping products, then validate counting depth on the physical prototype. The videos show the resulting local recognition workflow.

Mechanical Engineering

3D design · Arm retrofit · Camera installation

My Work

Camera control · Data pipeline · Models · Device integration

Why YOLO for This Prototype?

R-CNN was also tried. For a continuous camera feed, the selection focused on recognition under overlap, training iteration, and local deployment. YOLO’s single-stage detection pipeline suited this workflow; YOLOv7 with PyTorch was retained after comparison with YOLOv4.

Model Comparison

A Different Model. A Stronger Result.

Same dataset · Same EC2 · Equal epoch count

YOLOv7 · PyTorch≈99%

Final model · Test accuracy

>2×YOLOv7 / PyTorch vs. YOLOv4 training speed

Approximate historical test accuracy. Training speed compares the same data, image size, epoch count, and EC2 configuration. Model and framework changed together.

Implementation Workflow

  1. Device & Local Scripts

    Data Collection
    1. 01
      Position & Capture

      Move the camera to the target column’s center; capture images with Python / OpenCV.

    2. 02
      Prepare Images

      Organize images and labels for object-detection experiments.

  2. AWS Cloud Platform

    Cloud Training
    1. 03
      Store & Connect

      Amazon S3 stores the images; an Amazon EC2 GPU instance runs training on Linux. SSH keys provide secure remote access.

    2. 04
      Train & Select

      YOLOv4 with TensorFlow for earlier experiments; YOLOv7 with PyTorch for the final model.

  3. Physical Validation

    Local Deployment & Limits
    1. 05
      Integrate the Model

      Move the selected model from Ubuntu to Windows and connect the camera and positioning system.

    2. 06
      Filter & Count

      Use linear regression to help select center-column detections, then count the retained products and inspect the result on the prototype.

Cloud compute supported training; inference was integrated into the local prototype. Validation covers both recognition and the physical conditions needed for it.

DETECTION → COLUMN FILTER → COUNT

Count the Target Column

A camera frame can include products in neighboring columns. After detection, linear regression supports the column-selection step, filtering the results to the center column before counting. This keeps nearby products from being included in the target count.

YOLO · Find productsLinear Regression · Select the column
Center column · CountedAdjacent columns · ExcludedSelection concept · Not an experimental output

The Hard Part: Seeing Behind the Front Item

Products share the same line of sight. The task is to separate partially visible items, rather than count eight clearly separated objects.

PARTIAL OCCLUSION · SCHEMATICEdges remain visible on alternating sides
01020304050607–08 · HiddenREAR / LESS LIGHTCamera at the column center
Small left–right offsets expose different edges and labels. In the tested setup, the first six items were recognizable; the rear two had almost no usable visible area.
What the Model Must Learn

Use the remaining visible edges and appearance to distinguish overlapping products. Small exposed areas and dim light make the rear items harder.

Where the Limit Appears

Almost no visible pixels remain for the last two items. This identifies where lighting or viewpoint changes may be needed.

Prototype demonstration · 13 sec · silent

Company-owned work. Demonstration videos show the prototype output; source code, internal datasets, and non-public implementation details are not shared.

↑ Project index
03

GRAPHS & PREDICTION

NYCTraffic Prediction

Team course project · 2022

Connected Roads. Better Context.

Combining Uber Movement hourly speeds with OpenStreetMap, the model uses each road’s history and connected roads to predict the next hour’s speed.

Graph attentionPyTorchTime seriesExploratory analysis
New York City road network used in the traffic projectNew York City road network · Project visualization
Graph nodes
21,947
Graph edges
74,783
Temporal Inputs per Node
24

Why Look Beyond One Road’s History?

A road’s past speed describes its own rhythm; its connections show where neighboring traffic may matter. Graph theory represents those connections, and GAT learns how much neighbor information to use. Each node receives 24 hourly speed values: one measured variable across 24 time steps. The study also compared Prophet and Random Forest, testing the value of a network-aware approach.

Betweenness centrality distributions across congestion groups

Network Structure Offers Context

This comparison groups roads by congestion and examines how often each lies on shortest routes. The overlap shows that structure alone cannot explain traffic conditions; hourly speed histories are needed too.

MODEL OUTPUT

Predicted vs. Observed Speed

Red shows predictions and blue shows observations. Their broad rises, falls, and value ranges are similar, with several peaks and troughs occurring around the same time. The model captures much of this node’s overall pattern.

The prediction is smoother, and some sharp peaks and sudden drops remain imperfect. Look at both whether it follows the trend and where it misses abrupt changes.

One road-network node during the test period · Red: predicted · Blue: observed
One road-network node during the test period · Red: predicted · Blue: observed
INTERACTIVE SIMULATION

What Changes a Road’s Forecast?

Choose a scenario and a road. Edit an hourly speed or a connected road, then see how the next-hour estimate changes.

Start with an example
SelectedDirect neighbor

24-hour mean
Direct connections
Road D

Each hour is one input value. Changing an earlier hour changes the history without necessarily changing the latest speed.

60300−23 hNow+1 h
Blue: entered historyGold: next-hour estimate
Edit a Neighbor’s Latest Speed

Only directly connected roads contribute to this local information exchange.

mph

What Did Neighbor Information Change?
Own history only
With neighbors

Which Neighbors Get More Weight?

The project uses 24 hourly speeds per node and road connections to predict the next hour. This simulation makes those inputs editable with example data and illustrative calculations. The two estimates above explain the demo’s information flow, not a model-comparison score. Congestion cutoffs follow the study: 15.91 / 20.88 / 33.80 mph.

MIT License · Team: Harry Zhao, Yi Pan, Yantian Ding. Static figures and experiment metrics come from the team report; this was a course project, not a live traffic service.

↑ Project index
04

DATA & APPLICATIONS

StarTreeConfiguration Recommendations

Software Engineering Internship · May–Aug 2024

From Operating Load to Resource Advice.

Combine historical operating metrics and web data to support data-infrastructure configuration recommendations, then make the results available through a prediction API.

Pinot · KafkaSQL · Wget · GrafanaScikit-learn · FlaskREST · DockerPrompt Engineering
01

Data Preparation

SQL

Historical operating metrics

Kafka → Apache Pinot
Wget

Web data

Retrieve reference inputs

Clean, align and combine inputs

Grafana · SQL queries & historical load visualization

02

Prediction Service

Scikit-learn

Random Forest Regression

FlaskREST API

Make predictions available to the application

03

Configuration Advice

RECOMMENDATIONData Infrastructure

A configuration plan for the workload

Application Interface

Submit requirements and display the recommendations

DELIVERYDocker·GitHub ActionsContainerization & delivery workflow

The recommendation service turns historical data into a starting point for infrastructure planning. Its value is helping compare resources with workload needs, rather than presenting an isolated hardware number.

COMPONENTS & CONFIGURATION

Workload In. Configuration Out.

BrokerQuery routing

Read QPS · Query complexity · Latency target

Instance count · CPU · Heap / memory · Query concurrency

ServerData & computation

Read/write load · Data volume · Retention

Instance size · Memory · Disk capacity · Replication

ControllerCluster management

Table / segment count · Management load

Instance count · CPU / heap · Availability layout

ZooKeeperMetadata & coordination

Metadata size · Write load · Availability

Ensemble size · Heap · Log / snapshot storage

MinionBackground processing

Batch size · Task types · Completion window

Worker count · CPU / memory · Task concurrency · Schedule

Technical reference: StarTree Capacity Planning ↗ · Apache Pinot ↗

SQL + Grafana

Queried historical metrics with SQL in Grafana and visualized CPU usage to understand how resource load changed over time. This observation work provided context for the configuration-recommendation task.

INTERNAL AI EXPLORATION

Prompt Engineering

Configuration scenario · Illustration

Designed prompts and iterated on responses for an internal AI task. This example turns business requirements into a structured response with explicit fields, units, types and allowed values.

QPS
Query demand
Ingestion
Incoming data rate
Data & Constraints
Volume, retention, existing resources

Separate workload targets from existing infrastructure constraints.

JSONNumeric response · Example
{
  "configuration": {
    "broker": {"instances": 2, "vcpu": 4, "memory_gb": 16},
    "server": {"instances": 3, "vcpu": 8, "memory_gb": 32},
    "controller": {"instances": 2, "vcpu": 2, "memory_gb": 8},
    "zookeeper": {"instances": 3, "vcpu": 2, "memory_gb": 4},
    "minion": {"instances": 1, "vcpu": 4, "memory_gb": 16}
  }
}
JSON onlyConfiguration by componentTyped & bounded values
Output Contract
FieldTypeAllowed Values & Units
instancesinteger1–64ZooKeeper: 3 / 5 / 7; Minion may be 0
vcpuinteger2 / 4 / 8 / 16 / 32 / 64vCPU per instance
memory_gbinteger4 / 8 / 16 / 32 / 64 / 128 / 256GB per instance

Numbers illustrate the response format; the listed ranges define this demo interface, not cloud-provider limits or a plan for a specific workload. JSON ↗ · JSON Schema ↗

Company project. Internal code and data are not public.

↑ Project index

Technical PracticeBusiness Focus

Business needs guide the work. Technical evidence guides the decision.

BUILD & TEST

Understand What Can Be Built

Data analysis, model training, and application integration give me hands-on experience with how solutions work. I compare approaches through experiments and test what remains useful under real operating constraints.

Model TrainingData AnalysisIntegration
NEEDS & TRADE-OFFS

Connect It to a Business Need

I consider business goals, user needs, and cost constraints when defining a problem. I use data, technical feasibility, and practical feedback to compare options and inform product decisions.

User NeedsCost & DeliveryProduct Feedback

Understand the Business

Python · pandas · SQL

Use customer behavior, business processes, and data to identify problems, define useful measures, and support recommendations.

Test the Technology

PyTorch · TensorFlow · Scikit-learn · OpenCV

Compare models and implementations through experiments, drawing on computer vision, graph models, and Transformers.

Inform the Product

Flask · Git · Docker · AWS S3 / EC2

Connect models to real workflows and test their limits to inform features, retrofit options, and resource choices.

Contact

  • AI Product Manager
  • Machine Learning Engineer (MLE)
  • Business & Supply-chain Analytics
  • Data Analytics

Interested in roles connecting technology with products and business needs.

International OpportunitiesStrong interest in overseas work and travel

English & ChineseComfortable working across teams