ヒストグラムツールの使い方
How to Use the Histogram Tool
数値データから、ヒストグラム(度数分布のグラフ)・度数分布表・基本統計量を作成し、結果を Excel/CSV に保存して読み込み直すまでの操作方法をまとめています。
This page explains how to create a histogram (a graph of the frequency distribution), a frequency table, and summary statistics from numeric data, and how to save the results to Excel/CSV and load them again.
① ヒストグラムとは ② 入力方法 ③ 設定項目 ④ 結果の読み方 ⑤ 計算方法 ⑥ Excel/CSVの保存と読み込み ⑦ 上限と注意
① What Is a Histogram? ② Entering Data ③ Settings ④ Reading the Results ⑤ How It Is Calculated ⑥ Excel/CSV Save and Import ⑦ Limits and Notes
① ヒストグラムとは
① What Is a Histogram?
データの範囲をいくつかの区間(階級)に区切り、各階級に入るデータの数(度数)を棒で表したグラフです。データが「どのあたりに集まっているか」「左右対称か」「ばらつきはどのくらいか」といった、分布の形を見るために使います。
A histogram divides the range of the data into intervals (classes) and shows the number of values in each class (the frequency) as a bar. It is used to see the shape of a distribution: where the data cluster, whether it is symmetric, and how widely it spreads.
Cpkツールとの違い:このツールは規格値(上限・下限)を使いません。データそのものの分布の形を見るためのものです。規格値を前提にした工程能力(Cp・Cpk など)を評価したいときは、ポータルの「Cpk計算ツール」を使ってください。
Difference from the Cpk tool: this tool does not use specification limits. It is for looking at the shape of the data itself. To evaluate process capability against specification limits (Cp, Cpk, etc.), use the "Process Capability (Cpk) Calculator" on the portal.
③ 設定項目
③ Settings
- 階級の決め方(既定は「自動」)
- 自動(スタージェスの公式):階級数を「1 + log₂ n」(n はデータ数)を切り上げた値にして、最小値から最大値までを等分します。
- 階級数を指定:1〜50の整数を指定し、最小値から最大値までを等分します(既定は10)。
- 階級幅を指定:幅(0より大きい数)を指定します。階級は幅の整数倍のきりのよい値から始まります。この指定で階級数が50を超える場合はエラーです。
- 縦軸:「度数」(データの個数)か「相対度数」(度数 ÷ データ数)を選びます(既定は「度数」)。
- 度数ラベル:「あり」にすると、各棒の上に縦軸の値を表示します。度数が0の階級には表示されません(既定は「なし」)。
- 正規分布カーブ:チェックすると、正規分布の曲線(データの平均と標準偏差から計算)を重ねて表示します。データの形が正規分布に近いかどうかの目安になります(既定はオフ)。
- 統計量の表示:チェックすると、グラフの上にデータ数・平均・標準偏差を1行で表示します(既定はオン)。
- グループ間の縦軸(複数グループのときだけ効きます)
- そろえる(比較用)(既定):全グラフの縦軸の範囲を同じにします。グループ間で棒の高さを比べられます。
- 各グラフで自動:各グラフの縦軸をそのグラフのデータに合わせます。棒の高さをグラフ間で比べることはできません(形の比較には向いています)。
- グラフタイトル:任意です。空欄の場合は既定名 "Histogram" が使われます。
- How to Set the Classes (default: Automatic)
- Automatic (Sturges' formula): the number of classes is "1 + log₂ n" (n = number of values) rounded up, and the range from the minimum to the maximum is divided equally.
- Specify the number of classes: enter an integer from 1 to 50; the range from the minimum to the maximum is divided equally (default: 10).
- Specify the class width: enter a width (greater than 0). The classes start from a round multiple of the width. If this would need more than 50 classes, an error is shown.
- Vertical Axis: choose "Frequency" (the number of values) or "Relative frequency" (frequency ÷ number of values) (default: Frequency).
- Frequency Labels: when "On", the vertical-axis value is shown above each bar. Classes with a frequency of 0 get no label (default: Off).
- Normal Distribution Curve: when checked, a normal curve (calculated from the mean and standard deviation of the data) is overlaid. It is a guide to how close the data are to a normal distribution (default: off).
- Statistics on the Chart: when checked, the count, mean, and standard deviation are shown in one line above the chart (on by default).
- Y Axis Across Groups (applies only when there are several groups)
- Same scale (default): all charts share the same Y range, so bar heights can be compared between groups.
- Auto per chart: each chart's Y axis fits its own data. Bar heights cannot be compared between charts (it suits comparing shapes).
- Chart Title: optional. "Histogram" is used if left blank.
④ 結果の読み方
④ Reading the Results
- グラフ:横軸が階級の境界、縦軸が度数(または相対度数)です。複数グループのときは、全グループで共通の階級を使い、グループごとにグラフを縦に並べます。グラフ内の文字(統計量の行や凡例など)は英語です。
- 度数分布表:階級ごとの度数・相対度数(小数第4位まで)・累積度数・累積相対度数と、最下行の合計を表示します。画面では保存欄の下にあり、全グループの合計が31階級以上のときは閉じた状態で表示されます(見出しをクリックすると開きます)。
- 基本統計量:データ数・平均・標準偏差・最小・第1四分位数(Q1)・中央値・第3四分位数(Q3)・最大・歪度・尖度を表示します。複数グループのときは、グループごとに列が増えます。「—」は計算できない項目です。
- 表示桁数の違い:グラフ内の統計量は有効数字4桁、表は6桁前後で表示します。そのため、同じ平均でも 20.19(グラフ)と 20.1875(表)のように違って見えます(値は同じで、丸め方が違うだけです)。
- 入力の補正(見出しの読み飛ばし、重複した名前の連番など)があった場合は、結果の上に青い注記が表示されます。
- Chart: the horizontal axis shows the class boundaries and the vertical axis shows the frequency (or relative frequency). With several groups, all groups share the same classes and the charts are stacked vertically. Text inside the chart (the statistics line, legend, etc.) is in English.
- Frequency table: shows the frequency, relative frequency (to 4 decimal places), cumulative frequency, and cumulative relative frequency for each class, with a total in the last row. On the page it is below the save section; when all groups together have 31 or more classes, it starts collapsed (click its heading to open it).
- Summary statistics: shows the count, mean, standard deviation, minimum, first quartile (Q1), median, third quartile (Q3), maximum, skewness, and kurtosis. With several groups, there is one column per group. "—" means the value cannot be calculated.
- Different numbers of digits: statistics inside the chart use 4 significant digits, and the tables use about 6. So the same mean can look different, such as 20.19 (chart) and 20.1875 (table). The values are the same; only the rounding differs.
- If the input was adjusted (a header skipped, a sequence number added to a duplicate name, etc.), a blue notice is shown above the results.
⑤ 計算方法
⑤ How It Is Calculated
- 階級の作り方:「自動」と「階級数を指定」は、最小値から始めて最大値までを等分します。「階級幅を指定」は、最小値以下で幅の整数倍のうち最大の値から始めます(例:幅5で最小値が12なら10から)。
- 境界の扱い:階級は「以上〜未満」の [a, b) で、境界ちょうどの値は上の階級に入ります。最後の階級だけ最大値を含めるため [a, b](両閉)です。Excelの分析ツールやRの既定は「(a, b]」(境界の値が下の階級に入る)なので、境界上のデータがあると度数がずれることがあります。
- 複数グループ:全グループをまとめた範囲で、共通の階級を作ります。スタージェスの公式の n は、最大のグループのデータ数です。
- 統計量の定義(Excelの関数と同じ)
- 標準偏差は n−1 で割る標本標準偏差(STDEV.S)
- 四分位数は QUARTILE.INC と同じ(線形補間)
- 歪度は SKEW、尖度は KURT(正規分布で0になる超過尖度)と同じ定義
- 標準偏差は2件以上、歪度は3件以上、尖度は4件以上で計算できます。すべて同じ値のときの歪度・尖度も含め、計算できないときは「—」です。
- 正規分布カーブ:標本の平均と標準偏差で描きます。度数表示のときの高さは「データ数 × 階級幅 × 密度」、相対度数表示のときは「階級幅 × 密度」です(密度は正規分布の確率密度)。標準偏差が0または計算できないグループには描きません。
- How classes are made: "Automatic" and "Specify the number of classes" divide the range equally from the minimum to the maximum. "Specify the class width" starts from the largest multiple of the width that is at or below the minimum (e.g., width 5 with a minimum of 12 starts from 10).
- Boundaries: each class is [a, b) (includes a, excludes b), so a value exactly on a boundary goes into the upper class. Only the last class is [a, b] (closed on both sides) so that it includes the maximum. The Excel Analysis ToolPak and R default to "(a, b]" (a boundary value goes into the lower class), so the frequencies may differ when data lie exactly on boundaries.
- Several groups: common classes are made from the range of all groups combined. The n in Sturges' formula is the number of values in the largest group.
- Definitions of the statistics (same as the Excel functions)
- Standard deviation is the sample standard deviation, divided by n−1 (STDEV.S)
- Quartiles are the same as QUARTILE.INC (linear interpolation)
- Skewness is the same as SKEW, and kurtosis the same as KURT (excess kurtosis, which is 0 for a normal distribution)
- The standard deviation needs 2 or more values, skewness 3 or more, and kurtosis 4 or more. When a value cannot be calculated (including skewness and kurtosis when all values are identical), "—" is shown.
- Normal curve: drawn from the sample mean and standard deviation. In frequency display its height is "number of values × class width × density", and in relative frequency display it is "class width × density" (density is the normal probability density). It is not drawn for a group whose standard deviation is 0 or cannot be calculated.
例:データ A(16件)
Example: Data A (16 values)
10, 12, 14, 15, 15, 17, 18, 20, 20, 22, 23, 25, 25, 27, 30, 30
「自動」のとき:階級数 = 1 + log₂16 = 5、階級幅 = (30 − 10) ÷ 5 = 4 なので、境界は 10, 14, 18, 22, 26, 30 です。
With "Automatic": number of classes = 1 + log₂16 = 5 and class width = (30 − 10) ÷ 5 = 4, so the boundaries are 10, 14, 18, 22, 26, 30.
- 値 14 は境界ちょうどなので [14, 18) に入ります(14, 15, 15, 17 の4件)。値 30 は最後の階級 [26, 30] に含まれます。
- 基本統計量:平均 20.1875、標準偏差 6.18836、最小 10、Q1 15、中央値 20、Q3 25、最大 30、歪度 0.0938035、尖度 −0.981346。
- グラフの上の1行は「n = 16 Mean = 20.19 Std. Dev. = 6.188」(有効数字4桁)です。
- 「階級幅を指定」で幅を 4 にすると、最小値 10 ではなく 8(4 の倍数)から始まり、境界は 8, 12, 16, …, 32 になります。
- 正規分布カーブ(度数表示)の山の高さは、データ数 16 × 階級幅 4 ÷ (標準偏差 × √(2π)) ≒ 4.13 です。
- The value 14 lies exactly on a boundary, so it goes into [14, 18) (four values: 14, 15, 15, 17). The value 30 is included in the last class [26, 30].
- Summary statistics: mean 20.1875, standard deviation 6.18836, minimum 10, Q1 15, median 20, Q3 25, maximum 30, skewness 0.0938035, kurtosis −0.981346.
- The line above the chart reads "n = 16 Mean = 20.19 Std. Dev. = 6.188" (4 significant digits).
- If you specify a class width of 4, the classes start from 8 (a multiple of 4), not from the minimum 10, and the boundaries are 8, 12, 16, …, 32.
- The peak of the normal curve (frequency display) is 16 values × class width 4 ÷ (standard deviation × √(2π)) ≈ 4.13.
⑥ Excel/CSVの保存と読み込み
⑥ Saving and Importing Excel/CSV
保存:計算後の「結果を保存」から、グラフ画像(PNG)、Excel、CSV を保存できます。タイトルを入力している場合、グラフ画像(PNG)・Excel・CSVのファイル名はそのタイトル名になります(ファイル名に使えない記号は「_」に置き換わります)。空欄、または既定タイトル("Histogram")と同じ場合は、Excel・CSVは "histogram_result"、PNGは "histogram_chart" になります。
Saving: after calculating, use "Save Results" to save the chart image (PNG), Excel, or CSV. If a chart title is entered, it is used as the file name for the chart image (PNG), Excel, and CSV. (Characters not allowed in file names are replaced with "_".) If the title is blank or the same as the default title ("Histogram"), the files are named "histogram_result" (Excel/CSV) and "histogram_chart" (PNG).
- Excel(3シート):「Input Data」(設定のブロックと入力データ)、「Results」(度数分布表と基本統計量)、「Chart」(画像)。シート名・見出し・設定の値は英語で書かれます(以前の日本語の名前で出力したファイルも読み込めます)。
- CSV:設定のブロック、入力データ、度数分布表、基本統計量を1つのファイルにまとめます(画像は含まれません。UTF-8のBOM付き)。
- Excel (3 sheets): "Input Data" (the settings block and the input data), "Results" (the frequency table and summary statistics), and "Chart" (the image). Sheet names, headings, and setting values are written in English (files exported earlier with Japanese names can still be imported).
- CSV: the settings block, input data, frequency table, and summary statistics in one file (no image; UTF-8 with BOM).
読み込み:ページ上部の「Excel / CSVから読み込む」で、このツールが出力したファイルを選ぶと、設定・グループ名・データが復元され、結果が表示されます。
Importing: choose a file exported by this tool under "Import from Excel or CSV" at the top of the page, and the settings, group names, and data are restored and the results are shown.
- 入力方法は記録されません:方式Aで入力しても方式Bで入力しても、読み込むと「グループごとに入力」(方式A)の欄に復元されます。見た目と結果は同じで、変わるのは入力欄の形だけです。
- 設定のない単純なファイルも読み込めます:A列にデータを縦に並べた1列のファイル、または1行目に見出し(グループ名)を付けた複数列のファイル(最大5グループ)。この場合も、方式Aの欄に復元されます。
- このツールの出力ファイルなら、名前が数字(101 など)でも見出しとして読み込めます。
- CSVの ' について:CSVでは、名前が = + - @ で始まると、先頭に ' を付けて出力します(Excelなどで開いたときに数式として実行されないため)。Excelで開くと ' が見えますが、このツールで読み込むと外れて元の名前に戻ります。Excelファイルには ' は付きません(文字列として保存されます)。
- 1グループのときの見出し名は、図にも表にも出ないため、読み込み直すと保持されません(結果は同じです)。
- 重複した名前には「A (2)」のように連番が付きます。40文字の名前を重複させると、連番で40文字を超えるため、読み込みで名前の長さのエラーになります(名前を短くしてください)。
- CSV を Excel で開いて保存し直すと、Excel の仕様で、有効数字15桁を超える部分が失われることがあります。桁数の多いデータは、出力ファイルを Excel で上書き保存しないでください。
- ほかのツールの出力ファイルは読み込めません。
- The input method is not recorded: whether you entered the data with Method A or Method B, it is restored into the "Enter each group" (Method A) fields when imported. The appearance and results are the same; only the shape of the input fields changes.
- Plain files without settings can also be imported: a one-column file with the values down column A, or a multi-column file with group names (headers) in the first row (up to 5 groups). These are also restored into the Method A fields.
- For files exported by this tool, names that are numbers (such as 101) are read correctly as headers.
- About the ' in CSV: in CSV, a name that starts with = + - @ is written with a leading ' (so that it is not run as a formula when opened in Excel or similar). You will see the ' when you open the file in Excel, but it is removed and the original name is restored when you import it into this tool. Excel files have no ' (the names are stored as text).
- With one group, the header name appears in neither the chart nor the tables, so it is not kept when you import the file again (the results are the same).
- Duplicate names get a sequence number such as "A (2)". If a 40-character name is duplicated, the sequence number pushes it over 40 characters, and importing gives a name-length error (please shorten the name).
- If you open a CSV in Excel and save it again, digits beyond 15 significant digits may be lost (an Excel limitation). If your data has many digits, do not overwrite the exported files in Excel.
- Files exported by other tools cannot be imported.
⑦ 上限と注意:内容を変更したら、必ず「計算する」を押し直してください
⑦ Limits and Notes: After Any Change, Click "Calculate" Again
- 上限:データは全グループの合計で10,000件、グループは5つまで、階級数は50まで。グループ名は40文字まで、グラフタイトルは40文字までです。読み込むファイルは10MBまでです。
- このツールは自動更新されません。入力方法・グループ数・グループ名・データ・タイトル・階級の決め方・縦軸・グループ間の縦軸・度数ラベル・正規分布カーブ・統計量の表示のいずれを変更した場合も、「計算する」を押すまで結果には反映されません。データを直しただけの場合も、設定だけ変えた場合も同様です。
- Excel/CSVの保存は、最後に「計算する」を押した時点の入力と設定の内容で出力されます(入力欄を書き換えただけでは、保存の内容は変わりません)。
- データの取り扱い:入力・アップロードしたデータは、計算とグラフ作成のためにだけ処理され、サーバーに保存されません(通信はCloudflareを経由します。詳しくはプライバシーポリシーをご覧ください)。機密情報を含むデータは入力せず、品名・顧客名などは除くか記号に置き換えてご利用ください。
- Limits: 10,000 values in total across all groups, up to 5 groups, up to 50 classes, group names up to 40 characters, and the chart title up to 40 characters. The file you import can be up to 10MB.
- This tool does not update automatically. If you change the input method, number of groups, group names, data, title, class setting, vertical axis, Y axis across groups, frequency labels, normal curve, or statistics display, the results on screen will not update until you press "Calculate" again — whether you only edited the data or only changed a setting.
- Excel/CSV files are saved from the input and settings as of the last time you pressed "Calculate" (just editing the input fields does not change what is saved).
- Data handling: Data you enter or upload is processed only to perform calculations and create charts, and is not stored on the server (traffic passes through Cloudflare; see the Privacy Policy for details). Please do not enter confidential information; remove product or customer names or replace them with generic labels.