Ontology in Practice (2) : Four Tasks Behind a Sewing Data Dictionary

Ontology in Practice series 2/4 · (1) The Meaning of Data That Should Survive a System Change · (2) Four Tasks Behind a Sewing Data Dictionary · (3) Finding Meaning in a Database Without Keys · (4) Letting AI Answer on Top of a Meaning Layer
KEY POINTS
The sewing data dictionary is a data structure that gives processes, machines, styles and garment parts a unique ID, properties and relations. Through four tasks (normalizing process names, mapping machines, giving styles coordinates and attaching statistics to each entry), the process names in 1,021,940 process records were organized into 47,359 standard processes. Notations that cannot be judged are not merged on a guess but kept on hold.
22,287 operation breakdowns, and 1,021,940 process records written in them. That is the size of the operation breakdown records written between June 2022 and August 2026.1 When we first opened these records, the first question we ran into was this.
“Is this process the same process as that one?”
The process names in an operation breakdown are typed in by the analyst. Some people write the same process in capitals, some use abbreviations, some mix English with the local language, and machine names are no different. When the names differ, a computer treats the two records as different processes, so even collecting the standard times of one process becomes impossible.
The short answer is that SIJE calls the rules for answering this question, together with what those rules produce, the sewing data dictionary. The dictionary applies the components of an ontology covered in part 1 (class, instance, property and relation) to sewing process data. This article walks through the four tasks behind the dictionary from a developer’s point of view.
A glossary versus a data dictionary
SIJE already had a sewing process glossary. The glossary is a table that gathers the many names of the same process (Korean, English, Indonesian and Vietnamese) under one headword, and people use it for reading and translating.
A data dictionary serves a different purpose. So that a computer can merge and compare records, it gives every object a unique ID and records the values that object has and its relations to other objects alongside it.
A national ID system is a good comparison. One person may be called by a nickname at home, by a job title at work and by yet another name among friends, but there is only one ID number, and information such as address and date of birth is attached to that number. If the glossary is a notebook that collects all the names, the data dictionary is the registration system that lets you reach the same person through one number, however many names there are.
| Item | Sewing process glossary | Sewing data dictionary |
|---|---|---|
| Who reads it | People | Computers (and people) |
| Basic unit | Headword and its names in several languages | Unique ID (e.g. process_id) |
| What it holds | Different names with the same meaning | Name, property values, relations to other objects |
| Where it is used | Translation, training, communication on the floor | Merging records, rolling up standard times, finding similar styles |
The glossary and the data dictionary are different resources, but they are connected. In the first task of building the data dictionary, the glossary is the reference for deciding which names are synonyms.
Four components of the dictionary
The four components from part 1 map onto the sewing data dictionary as follows.
| Component | Definition | Sewing data dictionary example |
|---|---|---|
| Class | A category of things of the same kind | Standard process, machine, garment part, style |
| Instance | One object that belongs to a class | One standard process with a particular process_id |
| Property | A value attached to an object | Process name (full name), median SMV, coefficient of variation, sample count |
| Relation | A connection between two objects | Which garment part a process belongs to, which machine it is done on |
When stored as data, the relations are written in the RDF triple format from part 1 (subject, predicate, object).2
(Standard process #P) --belongs to--> (Part: hem)
(Standard process #P) --uses--> (Machine: lockstitch)
(Style #S) --contains--> (Standard process #P)※ #P and #S are placeholder numbers for explanation.
Once relations are written as triples, you can search from either end. You can start from a process and find the styles that used it, or start from a garment part and pull out every process that belongs to it.

Task 1. Normalizing process names
The first task is to group freely typed process names that describe the same process and give them one process_id. Normalization means turning values that are written differently but mean the same thing into one standard form.
The differences that show up in process names are things like capitalization, word order, synonyms and mixed languages. After these differences are cleaned up, notations that end up in the same form are grouped as candidates for the same process, and synonyms are judged against the sewing process glossary.
Input A: "Hemming Bottom Blind Stitch"
Input B: "BLIND STITCH BOTTOM HEMMING"
After cleaning up capitalization and word order, both take the same form
→ grouped as candidates for the same process※ Input A is an example for explanation.
Notations that mix languages are judged by checking them against the original process name. For example, the Indonesian word som means blind stitch, so notations in the som family are organized under the blind stitch process after the original process name has been checked.
Two processes are not merged just because their names look alike. Only notations that are clearly the same are grouped straight away, and similar candidates are merged conservatively. Through this process, the process names in 1,021,940 process records were organized into 47,359 standard processes.1

Task 2. Mapping machines
Machine names are managed in three tiers: the raw notation, the normalized form and the controlled vocabulary. The raw notation is exactly what the analyst wrote, the normalized form is that notation with its inconsistencies cleaned up, and the controlled vocabulary is the final list of machine codes the system accepts.
The hardest call here is deciding whether two notations are different names for the same machine or genuinely different machines. Merge them because the names look alike, and the times of different machines get mixed. Split them because they differ slightly, and the times of the same machine get scattered.
SIJE makes this call based on how the names are actually used, not on what they look like. If different people write the same machine in different ways, there is no reason for both notations to appear in an operation breakdown written by one analyst. On the other hand, if two notations always appear together in the same operation breakdown, they are being used as different machines within that breakdown.
For example, S/N and S/N#C look like the same family by name, but they always appeared together in the same operation breakdown, so they were split into separate machines. Iron and IRON, which differ only in capitalization, never appeared in the same operation breakdown, so they were merged as the same machine.

Task 3. Twelve style axes and garment part roles
Once processes and machines are organized, styles come next. SIJE describes the look and construction of a style with a controlled vocabulary on twelve axes. The axes include Body, Neckline, Sleeve, Placket, Pocket and Detail 1~4.1
The twelve axes play the role of latitude and longitude on a map. When every style has coordinates on the same twelve axes, a new style can be matched to past styles with nearby coordinates.
Notes on how people actually fill in the axes are kept with the vocabulary as well. For example, the Body axis for bottoms is meant to classify length, but on the floor people often write ‘Normal’ out of habit. If you classify without knowing these habits, the same style ends up with different coordinates.
Each process is assigned to one of 18 garment part roles.1 For a shirt, each item has a skeleton of parts such as collar, front, back, sleeve, cuff, hem, finishing and inspection, and each process is linked by relation to the part it belongs to.
Task 4. Attaching properties to each entry
Finally, statistical properties are attached to each standard process. For every entry, the median SMV, the coefficient of variation and the sample count are rolled up from the records.1
The median is the value in the middle when all values are lined up by size. Sewing process times sometimes include values stretched out to one side by interruptions or rework, so the median is a more stable representative value than the mean.
The coefficient of variation (CV) shows how spread out values are around the mean, as a ratio of the mean.
Coefficient of variation (CV) = standard deviation ÷ mean
Example) If a process has a mean SMV of 0.50 minutes and a standard deviation of 0.05 minutes,
CV = 0.05 ÷ 0.50 = 0.10※ The values above are an example to show how the calculation works.
A process with a small CV gives similar times across factories and analysts, so its median is easy to use as a reference. A process with a large CV varies a lot with fabric or detail conditions, so it has to be broken down by condition. The sample count tells you how many records a median came from, so you can judge how much evidence stands behind two medians that look the same.
Rules we kept while building the dictionary
Three rules were kept throughout the four tasks.3
| Rule | What it means | Why |
|---|---|---|
| Hold what cannot be judged | Notations without enough evidence are not merged on a guess but kept on hold | A wrongly merged process is far harder to find and split later |
| Keep the history of rejections | Hypotheses rejected in review are not deleted but kept as history | It prevents the same hypothesis from being reviewed again and keeps the basis for each judgment traceable |
| A universal dictionary, independent of any factory | The dictionary holds only definitions that hold in any factory, while the physical attributes of a specific line live in the factory and line master | If one factory’s conditions get mixed into the dictionary, other factories cannot use it |
The reason for the third rule is simple. If equipment conditions, such as what a particular line can measure, are written into the definition of a process in the dictionary, then in a factory without that line the same process will look like a different process.
Where the dictionary is used
Once the dictionary is complete, every measured value can be given semantic coordinates. Each value SIJE registers as an asset is recorded together with which standard process it is the time of, and which garment part, machine and style coordinates it corresponds to. The data dictionary is what makes that connection.3
A value without semantic coordinates cannot tell you which process it is the time of, so it cannot be reused on the next order. A value with semantic coordinates, on the other hand, can be grouped with records from other factories and other periods for the same process and compared.
A typical feature that depends on the dictionary is drafting an operation breakdown. When a new style comes in, similar past records can be found based on its twelve axis coordinates and garment part roles, and a draft operation breakdown can be built from them. This works because the dictionary knows whether processes from different styles are the same standard process, and which garment part each process belongs to.
In SIJE’s six stages of turning manufacturing knowledge into assets, this dictionary handles stage 2 (giving data meaning).3 Cycle records collected by Monolog in stage 1 are given meaning with this dictionary, and in stage 3 they are filtered with statistical algorithms to set the reference value for each process. The next part applies the same principles to the sales data of a fashion brand instead of sewing processes, and looks at how to find meaning in a database that has no linking information.

Sewing Production Glossary
Normalization
Turning values that are written differently but mean the same thing into one standard form.
Controlled vocabulary
A closed list of standard terms the system accepts. Notations that are not on the list cannot be entered.
Triple
A data format that records a relation in three slots: subject, predicate and object.
process_id
The unique ID given to one standard process in the standard process dictionary.
SMV (Standard Minute Value)
The standard work time, in minutes, a skilled operator needs to perform a process once using the specified method.
Coefficient of variation (CV)
The standard deviation divided by the mean, showing how spread out values are as a ratio.
👉 Ontology in Practice (1) : The Meaning of Data That Should Survive a System Change
👉 Ontology in Practice (3) : Finding Meaning in a Database Without Keys
References
- SIJE internal data. Size and period of the operation breakdown records, and the structure of the standard process dictionary and style classification. ↩
- W3C, RDF 1.1 Concepts and Abstract Syntax, 2014. Definition of the triple made of subject, predicate and object. ↩
- SIJE internal data. Stage definitions of the six stages of turning manufacturing knowledge into assets, and the design principles of the data dictionary. ↩
