Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Automated Wafer Inspection

Two complementary models for wafer inspection, both served by one web app (python -m webapp.server, see Web app):

  1. Bin-map pattern classifier — a CNN classifying the spatial pattern failing dies form on a wafer map (Center, Scratch, Edge-Ring, etc.), trained on WM-811K (LSWMD).
  2. Surface defect detector — a YOLOv8-OBB model that finds and classifies real physical defects (particle contamination, etching problems, etc.) anywhere in a wafer surface photo, trained on a real fab defect-photo dataset.

These answer different questions: (1) is "what shape do the failing dies form," (2) is "what real physical defect is present, and where." A pattern classifier alone can't tell you why dies failed; a bin map alone doesn't localize physical defects.

Setup

python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt

# The surface-defect model (below) needs a CUDA-enabled torch build for GPU training;
# requirements.txt alone installs CPU-only torch. For GPU:
pip install torch==2.13.0 torchvision==0.28.0 --index-url https://download.pytorch.org/whl/cu130

1. Bin-map pattern classifier

Defect classes: Center, Donut, Edge-Loc, Edge-Ring, Loc, Near-full, Random, Scratch, none.

A note on the raw dataset

dataset/raw/LSWMD.pkl was pickled with pandas 0.14.1 under Python 2 (2018-era). Modern pandas (3.x, installed here) cannot unpickle it directly, and no legacy pandas/numpy install is needed to fix it — src/data/convert_legacy_pickle.py walks the raw pickle opcodes with a custom unpickler that treats every pandas internal class as a no-op placeholder, since pickle only needs opcodes to execute, not for reconstructed objects to be "real" pandas objects. See that file's docstring for details if it ever needs to be revisited.

Pipeline

Run each step from the project root, in order:

# 1. One-time: convert the legacy pickle into modern, version-agnostic files
#    -> dataset/processed/{images.npy,labels.csv,label_classes.json}
#    -> dataset/{train,validation,test}/indices.npy (stratified 70/15/15 split)
python -m src.data.convert_legacy_pickle

# 2. Train the CNN (checkpoints -> models/checkpoints/best_model.keras,
#    history plot -> results/plots/training_history.png)
python -m src.train --epochs 30 --batch-size 128

# 3. Evaluate on the held-out test split
#    -> results/metrics/classification_report.json
#    -> results/confusion_matrix/confusion_matrix_{raw,normalized}.png
#    -> results/predictions/sample_predictions.png
python -m src.evaluate

Dataset notes

Of the 811,457 wafers in LSWMD, only 172,950 have a defect-type label; the rest are unlabeled and are not used here. The labeled subset is heavily imbalanced (~85% are defect-free none), handled during training via class weighting (sklearn.utils.class_weight.compute_class_weight) and orientation augmentation (random flips/90° rotations — wafer maps have no fixed orientation).

2. Surface defect detector

Real physical defect classes: BLOCK ETCH, COATING BAD, PARTICLE, PIQ PARTICLE, PO CONTAMINATION, SCRATCH, SEZ BURNT.

Dataset

Wafer Defect by Roboflow user wafer-irhuv, licensed CC BY 4.0 (attribution required if redistributed) — 4,532 real wafer surface photos with YOLOv8-OBB (oriented bounding box) labels. Not committed to this repo (see .gitignore); to re-obtain it, download the YOLOv8 export from that URL and extract to dataset/surface_defect/.

About ~22% of images are skipped automatically by Ultralytics at load time ("ignoring corrupt image/label: non-normalized or out of bounds coordinates") — a quirk of this specific export where some rotated-box corners fall slightly outside [0,1] after Roboflow's rotation math. This is expected and not a bug in this repo; training proceeds on the remaining ~78%.

Pipeline

# Train (checkpoints -> models/surface_defect/train/weights/best.pt)
python -m src.surface_defect.train --epochs 100 --imgsz 640 --batch 16

# Evaluate on the held-out test split
#    -> results/surface_defect/metrics.json
#    -> results/surface_defect/sample_predictions/
python -m src.surface_defect.evaluate

Trains on GPU (device=0) via PyTorch/Ultralytics — unlike TensorFlow, which has no native-Windows GPU support for TF≥2.11 (the bin-map CNN above trains on CPU for that reason). On an RTX 4060, each epoch takes roughly 30 seconds.

Web app (upload an image, get a full inspection report)

python -m webapp.server        # then open http://127.0.0.1:8000

or double-click Open Wafer Web App.bat. Starlette + uvicorn, both already in .venv -- no extra dependency. Set PORT to run it somewhere else (PORT=8010 python -m webapp.server).

Drop in a PNG/JPG of a wafer map, a .npy bin map, or a photo of a wafer surface. One scan runs both models and reports both:

  • the bin-map pattern CNN, which names the spatial pattern the failing dies form, and
  • the YOLOv8-OBB surface detector, which finds physical defects in a photograph.

Which one an upload suits is still worked out, but only to rank the two: the matching model is marked best match and the other out of domain, with a caveat banner saying why its reading is unreliable, rather than being discarded. Routing to one model and throwing the other away means a wrong guess costs the whole result; running both costs about a second and never does. Either section can also report that it could not run -- a .npy array has no photograph for the detector, and a photograph often has no recoverable die grid. Severity is only scored when the bin-map reading is in domain.

The report then covers:

  • Classification -- the WM-811K pattern class with its full probability distribution, plus that class's real precision/recall/F1 from the last evaluation run, so a Donut call (F1 0.43) is not presented with the same confidence as an Edge-Ring call (F1 0.97).
  • Where the damage is -- failing dies are clustered (8-connected), and each cluster is placed in wafer coordinates: radial zone (centre / mid-radius / edge), clock position, compass sector, radius as a fraction of the wafer radius, and shape (a Scratch reports its length and bearing, a ring reports how much of the circumference it covers). Backed by the fail-rate-per-annulus chart, so "edge ring" is checkable rather than asserted.
  • Wafer profile -- die grid, dies on wafer, failing dies, yield, failure rate, map density band.
  • What the pattern means at the stage you inspected it (see below) -- the root causes that stage allows, the ones it rules out, a suggested next step, and whether the wafer is still reworkable.
  • Model attention -- a Grad-CAM overlay plus a check on whether the heat actually sits on the failing dies. When it does not, the call is flagged as suspect however confident the softmax looks.

Two model behaviours are worth knowing about:

  • Inference uses 8-way dihedral TTA. Wafer maps have no canonical orientation (training augments with flips and 90-degree rotations), so the softmax is averaged over all 8 rotations/reflections in one batched forward pass. Measured on 6,000 held-out test wafers: accuracy 0.9620 -> 0.9643, macro-F1 0.7612 -> 0.7646.
  • Position is measured, not predicted. In a bin map the failing dies are marked explicitly, so src/inspection/geometry.py computes the location directly instead of asking a model to guess it. The CNN is only asked what it is good at: naming the pattern.

Stage-aware diagnosis

The same signature means different things depending on when it is caught. An edge ring found at develop inspect points at edge bead removal on the track; the identical ring found after etch points at focus-ring erosion, and EBR is no longer worth chasing -- if the resist edge were wrong, the previous inspection would have caught it.

So pick the inspection point in the sidebar (after deposition, ADI, AEI, post-CMP, or wafer sort) and the report changes accordingly:

  • every cause is tagged with the process module that produces it, and modules that have not run yet at that stage are moved into an explicit ruled out list -- often the more useful half, since it says where not to look;
  • handling damage is never excluded, because wafers are moved and chucked at every step;
  • the disposition changes with the stage: resist can be stripped and re-coated at ADI, while after etch the damage is permanent.

An Edge-Ring call, same wafer, three stages:

Inspected at Causes offered Ruled out Disposition
After develop (ADI) EBR width, film roll-off plasma etch, probe map reworkable
After etch (AEI) + plasma etch / focus ring probe map not reworkable
Wafer sort all four none end of line

Two honest limits, both stated in src/inspection/stages.py. This models one layer of a simplified flow (deposit, pattern, etch, clean, planarise, test), where a real device is hundreds of such loops. And WM-811K maps are electrical bin maps from wafer sort -- the last stage in the list -- so choosing an earlier stage changes the interpretation and never the classification; earlier stages describe optical defect maps from in-line inspection, which are the same shape of data but not what the classifier was trained on.

Reading arbitrary images back into a die grid

src/inspection/ingest.py recovers a bin map from any rendering. Images in this project's own palette are matched exactly; anything else is segmented (page margin trimmed, corner-filling colours dropped as off-wafer, k-means inside the disk to split pass from fail, warmer or brighter cluster taken as failing -- use the "Swap pass / fail colours" toggle if that guess comes out backwards). The die grid itself is recovered from the fact that every colour transition in a rendered map lands on a die boundary.

Round-trip checked: a known WM-811K map rendered in five styles (three colour schemes, three sizes) comes back with the correct 64x64 die grid in all five and a failure rate within ~1 percentage point.

Uploads that are photographs rather than maps are detected by colour flatness (top-6 colours cover 90-99% of a rendered map against 25-53% of a real photo) and routed to the surface-defect detector instead.

Streamlit app

streamlit run app/streamlit_app.py

Three modes in the sidebar: bin-map demo (random/indexed test-set sample), bin-map .npy upload, and surface-defect detection (demo test photo or your own upload).

If streamlit/python on your PATH don't resolve to this project's .venv (e.g. venv activation not taking effect in your shell), call the venv's Python by full path instead: .venv\Scripts\python.exe -m streamlit run app\streamlit_app.py.

Demo notebook

notebooks/demo.ipynb — a simpler, no-server alternative to the Streamlit app: run cells one at a time (python -m jupyter lab, then open the notebook) and see each model's prediction rendered inline, research-notebook style. Re-run either demo cell to try a new random sample.

Note: the jupyter-*.exe launcher scripts in .venv\Scripts are broken on at least one tested setup (exit 1, no output, even on --help). Use python -m jupyter lab / python -m nbconvert (module invocation) instead of the .exe launchers.

Tests

pytest tests/

Project layout

src/
  config.py                    shared paths/constants for both models
  data/
    convert_legacy_pickle.py   one-time legacy pickle -> processed bin-map dataset
    dataset.py                 tf.data loading, class weights (bin-map classifier)
  models/
    cnn.py                     CNN architecture (bin-map classifier)
  train.py, evaluate.py        bin-map classifier train/eval
  surface_defect/
    train.py, evaluate.py      YOLOv8-OBB train/eval (surface defect detector)
    inference.py               inference helper used by the apps
  inspection/                  serving side: upload -> full inspection report
    ingest.py                  any image/.npy -> canonical die grid, or route to the detector
    stages.py                  process-flow model: which causes each inspection stage allows
    geometry.py                clusters, radial/clock position, yield, radial profile
    predictor.py               checkpoint loading, TTA prediction, Grad-CAM
    knowledge.py               per-class fab meaning, root causes, severity
    render.py                  report figures as base64 PNGs
    pipeline.py                orchestrates the above into one report dict
webapp/
  server.py                     Starlette/uvicorn API + static hosting
  static/                       the page (index.html, style.css, app.js)
app/
  streamlit_app.py              older app: bin-map demo/upload + surface-defect detection
notebooks/
  demo.ipynb                    no-server cell-by-cell demo of both models
tests/
dataset/
  raw/                          original LSWMD.pkl (not committed)
  processed/                    converted bin-map images/labels (not committed)
  train/, validation/, test/    bin-map split indices (not committed)
  surface_defect/                Roboflow photos + YOLO-OBB labels (not committed)
models/
  checkpoints/                  bin-map CNN checkpoint (not committed)
  surface_defect/                YOLOv8-OBB run output/checkpoints (not committed)
results/
  plots/, confusion_matrix/, metrics/, predictions/   bin-map classifier results
  surface_defect/                                     detector metrics + sample predictions

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages