Explore a LAS curve’s distribution and compare selected samples with all included samples.
Danila Karnaukh · Petrophysics.io · v18 · 2026-10-04
Open Histogram · Python reference · Lock dock
Start with a LAS
- Open Histogram from the home page or any tool’s switcher.
- Choose or drop a well-log LAS file, or click Try with sample data. With Lock dock on, Histogram uses the shared original LAS from the other tools.
- Choose a Distribution curve. GR is selected when available; otherwise a numeric non-index curve is preferred. Curves retain their original units. Calculated Python curves appear in the same selector.
- Use Distribution to choose bins, X scale, Y display and optional analysis limits. Click Apply distribution. Controls take effect when applied.
- Hover over bars for exact bin intervals, counts, percentages and cumulative percentages. The tables below provide the same values without using the pointer.
- Read the eight headline statistics, then expand All statistics · compare selection for the full table.
- Download Statistics CSV, Bins CSV or Statistics JSON. Use Export PNG for an image and Save project to reproduce the view.
The sample is synthetic, not a real well. Processing is local to the browser. The first Python run needs internet access to download its runtime; ordinary histogram calculation needs no Python runtime.
What the numbers describe
Statistics use unbinned original values, with equal weight per source row. Changing the number of bins or Y display does not change the mean, median or percentiles. Log X changes which samples qualify but does not log-transform the reported statistics: a GR mean is still in API, for example.
The exclusion counts form a disjoint sequence, in this order:
- Missing/non-finite distribution-curve values, including the LAS NULL marker.
- Samples outside the inclusive depth/index limits, or with missing depth when those limits are active.
- Samples failing the curve filter, including missing filter-curve values.
- Non-positive values when X uses a logarithmic scale.
- Samples outside the inclusive analysis minimum/maximum.
- The remaining samples are included.
These counts sum to the source row count; a sample excluded at an earlier stage is not counted again. Selection is an overlapping subset, not another exclusion stage. The table compares All included with Selected included. Selected rows hidden by a filter or range stay selected but do not contribute until included again. Their count is shown separately.
Analysis limits and filters change the population. Zoom only changes the view. Both statistics and bins continue to describe the complete analysis while zoomed. Reset view / double-click clears zoom only. Clear analysis limits and Clear all filters are separate controls.
Bin methods and axes
Let N be the included sample count, R the sample range, IQR the interquartile range, and σ the population standard deviation. On log X these bin estimators operate on log10 values; statistics still use raw values.
| Method | Rule |
|---|---|
| Auto · FD / Sturges | Smaller positive width from FD and Sturges. If IQR is zero, use Sturges. This explicit rule is version-independent and may differ from another library’s auto. |
| Freedman–Diaconis (FD) | Width h = 2 × IQR / N^(1/3). IQR = 0 falls back to Sturges with a notice. |
| Scott | Width h = σ × (24 × √π / N)^(1/3). |
| Sturges | Target count log2(N) + 1; width R / (log2(N) + 1). |
| Square root | Target count √N; width R / √N. |
| Custom count | An integer from 1 to 500. |
| Custom width | A positive width in curve units on linear X, or log10 units on log X. The last bin may be narrower. |
Automatic methods compute ceil(bin range / h), then divide the bin range into equal-width bins in the chosen X scale. They are capped at 500 bins, with a notice. A custom width requiring more than 500 bins is rejected rather than silently changed. Zero-spread automatic data uses one bin; constant ranges are expanded for display (linear: ±max(5% of absolute value, 0.5); log: ±0.5 in log10 space). Explicit analysis bounds also bound the bins.
Bins include their lower edge and exclude their upper edge: [lower, upper). Only the last bin also includes its upper edge. Every included sample belongs to exactly one bin. An exact repeated-value mode and a highest-count bin are different concepts; the JSON report includes the indices of all highest-count bins.
| Y display | Meaning |
|---|---|
| Count | Number of included samples in each bin. |
| Percentage | 100 × bin count / included N. All-bin percentages sum to 100%. |
| Probability density | bin count / (included N × raw bin width). Sum of density × raw width is 1, including with log-spaced bins. Visual area on a log axis is not a probability. |
The optional cumulative line shows the percentage of all included samples at or below each bin’s upper boundary, with a separate 0–100% right axis. Orange selected bars use the same all-included denominator as the base histogram, so they show the selected contribution rather than a separately normalized distribution. Reversing X only reverses the display.
Selecting samples and zooming
- Default interaction is Select interval. Click a bar to select that bin or drag horizontally to select an inclusive value interval.
- In Selection, use Replace, Add or Remove. You can also enter minimum/maximum values and click Select value interval. Blank limits select through the included extremes.
- Expand Bin counts and cumulative distribution, focus a bin button and press Enter to select it with the current selection mode.
- For zoom, choose Interaction → Zoom view, then drag on the chart. Choose Reset view or double-click to show the complete bin range again.
- Selection uses original zero-based row IDs. It survives filters and bin changes. Choosing a different distribution curve clears selection and analysis limits; loading a different LAS resets the view and results.
- Expand / Dock is available for the whole Histogram panel and Python editor. Escape docks an expanded panel. The curve picker, filters and statistics remain available while Histogram is expanded.
Statistics and formulas
N denotes the included sample count for the relevant column (all or selected), x a raw value, and μ its arithmetic mean. A dash / JSON null means empty, undefined or outside floating-point range. Counts are 0 for an empty population. Most values have the curve’s original unit; variances use its square. Skewness, excess kurtosis and CV are dimensionless (CV is displayed as %).
| Statistic | Definition and requirements |
|---|---|
| Count, unique, positive, zero, negative | Number of samples, distinct numeric values and sign counts. Missing values are already excluded. |
| Minimum, maximum, range, sum | Extrema, max − min and sum of included values. |
| Mean, median | Arithmetic mean and P50. |
| Exact mode | Most frequently repeated value; smallest value if several tie. Undefined if all values occur once. Frequency and number of tied values are reported separately. This is not a smoothed density estimate. |
| Population variance / standard deviation | Σ(x − μ)² / N, and its square root. |
| Sample variance / standard deviation | Σ(x − μ)² / (N − 1), and its square root; N ≥ 2. |
| Standard error of mean | Sample standard deviation / √N; N ≥ 2. Interpretation assumes independent samples. Adjacent depth readings are often correlated, so this is not an automatically valid uncertainty estimate. |
| Coefficient of variation (%) | 100 × sample standard deviation / abs(μ); N ≥ 2, μ ≠ 0. Use carefully with near-zero means or units having an arbitrary zero. |
| RMS | √(mean(x²)). |
| Percentiles | P1, P5, P10, P25, P50, P75, P90, P95, P99. Sort values; linearly interpolate at index (N − 1) × p/100 (Hyndman–Fan type 7). |
| Q1, Q3, IQR | P25, P75 and Q3 − Q1. |
| MAD | Median(abs(x − median(x))), without a normal-distribution scale factor. |
| Skewness | Adjusted Fisher–Pearson: √(N(N−1))/(N−2) × m3/m2^(3/2), where mk = mean((x−μ)^k). Needs N ≥ 3 and nonzero variance. |
| Excess kurtosis | (N−1)/((N−2)(N−3)) × ((N+1)(m4/m2²−3)+6). Bias-corrected Fisher convention: normal-distribution reference is 0, not 3. Needs N ≥ 4 and nonzero variance. |
| Geometric mean | exp(mean(log(x))); all included values must be strictly positive. |
| Harmonic mean | N / Σ(1/x); all included values must be strictly positive. |
| Tukey fences / outlier count | Q1 − 1.5 × IQR and Q3 + 1.5 × IQR; count strictly outside those fences. These flags do not remove values or prove bad data. |
The CSV keeps numeric values at JavaScript’s full available precision; the UI rounds for readability. A constant curve has zero variance and undefined skewness/kurtosis. Single-sample sample variance/SD/SE/CV are undefined. Extremely large ranges or excessively close bin edges may be rejected with an explicit message; individual unrepresentable statistics are null.
Python, LAS export and reproducibility
Python has the same well interface as the other tools and also works without a LAS (well is None). New script starts with an empty editor. With a LAS, Python receives the full source rows, unaffected by histogram filters, selection or zoom. well.visible_interval is the full depth/index range here.
import numpy as np
if well is None:
print("Load a LAS to calculate a result curve.")
else:
gr = well["GR"]
# Example endpoints only: choose references appropriate to your well.
vsh = np.clip((gr - 20.0) / (120.0 - 20.0), 0.0, 1.0)
well.add_curve("VSH_PY", vsh, unit="v/v",
description="Linear GR index; example endpoints")
After a successful run, click Use in histogram in Python Results or choose VSH_PY in the curve selector. Source LAS curves are not overwritten. Export LAS includes all original and committed calculated curves over the full LAS interval, irrespective of filters, selection or zoom. See the Python reference for naming, package, worker and timeout limits.
| Download | Contents |
|---|---|
| Export PNG | Labelled current chart viewport with summary statistics and source fingerprint. |
| Statistics CSV | Full all/selected statistics, exclusion counts, curve/unit, source fingerprint, settings and definitions. |
| Bins CSV | Every bin’s edges, boundary convention, count, selected count, percentage, raw-unit density and cumulative percentage, plus analysis metadata. |
| Statistics JSON | Structured report with both populations, counts, settings, definitions and bins. This is a report, not a reloadable project. |
Save project (.scene.json) |
Histogram configuration, selection, viewport, Python script and result curves. Load the matching original LAS before Load project. SHA-256/row count prevent applying it to a different source. Loading never executes Python. |
| Export LAS | Complete parsed LAS, including committed Python curves. Not limited to selected samples. |
| Save .py | The Python script. Useful even with no LAS loaded. |
Lock dock shares the original source LAS and selected dataset, not scripts, derived curves, histogram settings or selections. To share a computed result, export LAS and load that exported file while locked. Closing/reloading a tool discards unsaved local view/script/results; save a project first.
JavaScript API
The in-page API is window.Petrophysics; no remote plotting HTTP API is added. Consult get_capabilities() for the current tool. Histogram supports all shared dataset/project/Python/Lock dock commands plus:
// Use after importing a LAS. Start from the complete current configuration.
const scene = await Petrophysics.get_scene();
scene.config.distribution.method = "count";
scene.config.distribution.count = 25;
scene.config.distribution.measure = "percent";
scene.config.viewport = null;
await Petrophysics.configure_histogram(scene.config);
// Select original source rows, not bin indices.
await Petrophysics.set_selection({rows: [0, 1, 2], mode: "replace"});
const report = await Petrophysics.get_histogram_statistics();
console.log(report.statistics.mean, report.selectedStatistics.mean);
console.table(report.bins);
await Petrophysics.export_result({format: "statistics-csv"});
await Petrophysics.export_result({format: "bins-csv"});
const result = await Petrophysics.export_result({format: "statistics-json", download: false});
console.log(await result.blob.text());
configure_histogram validates the complete configuration before changing state. Curve references have {id, mnemonic, unit}; filter comparisons use gt, gte, lt, lte, eq, neq. get_histogram_statistics() returns an independent report snapshot. Report bin indices are zero-based; CSV/UI bin numbers start at 1. Scene schema: scene-v1.json.
Limits and interpretation
One LAS at a time; existing import limits: 25 MiB, 250,000 rows, 5,000,000 data cells across all tables. Python result limits remain 64 curves / 5,000,000 output values. Bins are limited to 500. No multi-well overlay, fitted probability distributions, confidence intervals, hypothesis tests, thickness weighting or automatic geological interpretation is added. Irregularly spaced or repeated depth rows still get equal weight. Choose units and populations deliberately. This tool reads well-log LAS versions through 3.0, not LiDAR LAS.
Use HTTPS or a local HTTP server; double-clicking file:// pages does not support the Python worker/scene fingerprint workflow. Chat sends typed chat messages separately; the site assistant cannot automatically inspect your loaded histogram or execute code.
Formula references
- NumPy histogram bin estimators. The explicit Auto rule and fallbacks above are authoritative for this tool.
- NumPy percentile conventions.
- NIST skewness and kurtosis discussion. The bias corrections and excess convention used here are stated above.
LAS datasets in v22
All four tools share the reader for LAS versions through 3.0. Choose the active table using LAS dataset. Projects record its identity; Lock dock shares that choice and the source file. Numeric plots require numeric data, and Log Viewer requires depth. Python exposes typed fields and all tables with well.datasets and well.dataset(id). Legacy numeric export remains LAS 2.0; LAS 3.0 export retains extra tables and typed fields. LAS formats, syntax and limits.