> For the complete documentation index, see [llms.txt](https://stereotoolss-organization.gitbook.io/saw-user-manual-v8.1/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://stereotoolss-organization.gitbook.io/saw-user-manual-v8.1/analysis/outputs/html-report.md).

# HTML report

`SAW count` and `SAW realign` pipelines will output an interactive report `<SN>.report.html` . The contents of the HTML report file will vary depending on the pipeline and parameters used but generally follow a similar format across runs.

On this page, we demonstrate the report of&#x20;

* a mouse brain sample from Stereo-seq T FF V1.3,
* a mouse lung tissue sample from  Stereo-seq N FFPE V1.0,
* and a mouse thymus tissue sample from Stereo-CITE T V1.0.

Run with `SAW count` v8.1.

## **Summary**

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FtgKa1wP0V83BPIr24Ji5%2Fsummary-exprssion.PNG?alt=media&amp;token=a0d60c1e-5ee7-4ace-b691-ec72c93a8e69" alt=""><figcaption><p>Expression heatmap and four key metrics</p></figcaption></figure>

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FYcGI4WQrgafjhSYyPANp%2Fsummary-image.PNG?alt=media&amp;token=a41deadb-facb-497b-8aee-982672c954ee" alt=""><figcaption><p>Display of microscope image</p></figcaption></figure>

The spatial gene expression distribution plot, containing all bins and bins under tissue, on the left, shows MID count at each bin20.

`Total Reads` is the amount of total sequencing reads of input FASTQs. `Mean MID per Bin20/Mean MID per Bin50` and `Mean Gene per Bin20/Mean Gene per Bin50` represent the mean MID and gene type counts at each bin20 or bin50. `Total Genes` is the number of gene types from all bins.

### Key metrics

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FUM4YEvN1Aqb92hvQA1cq%2Fsummary-key-metrics.PNG?alt=media&amp;token=efcc4002-3fc9-4514-841f-5e5ab48974d1" alt=""><figcaption><p>Details and sunburst plot of key metrics</p></figcaption></figure>

Key metrics of the data are listed:

<table><thead><tr><th width="227">Metric</th><th>Description</th></tr></thead><tbody><tr><td>Total Reads</td><td>Total number of sequenced reads.</td></tr><tr><td><a href="/saw-user-manual-v8.1/algorithms/gene-expression-algorithms.md#cid-mapping">Valid CID Reads</a></td><td>Number of reads with CIDs that can be matched with the mask file</td></tr><tr><td>Invalid CID Reads</td><td>Number of reads with CIDs that cannot be matched with the mask file.</td></tr><tr><td>Clean Reads</td><td>Number of Valid CID Reads that have passed QC.</td></tr><tr><td><a href="/saw-user-manual-v8.1/algorithms/gene-expression-algorithms.md#rna-filtering">Non-Relevant Short Reads</a></td><td>Number of non-relevant short reads</td></tr><tr><td>Discarded MID Reads</td><td>Number of reads with MID that have been discarded since MID sequence quality does not satisfy with further analysis.</td></tr><tr><td>Uniquely Mapped Reads</td><td>Number of reads that mapped uniquely to the reference genome. If the pipeline uses uniquely mapped reads and the best match from multi-mapped reads for subsequent annotation, this item will include them both</td></tr><tr><td><a href="/saw-user-manual-v8.1/algorithms/gene-expression-algorithms.md#annotation-to-transcriptome">Transcriptome</a></td><td>Number of reads that are aligned to transcripts of at least one gene</td></tr><tr><td>Unique Reads</td><td>Number of reads in Transcriptome that have been corrected by MAPQ and deduplicated</td></tr><tr><td>Sequencing Saturation</td><td>Number of reads in Transcriptome that have been corrected by MAPQ with duplicated MID</td></tr><tr><td>Unannotated Reads</td><td>Number of reads that cannot be aligned to the transcript of one gene</td></tr><tr><td>Multi-Mapped Reads</td><td>Number of reads that mapped more than one time on the genome. If the pipeline uses uniquely mapped reads and the best match from multi-mapped reads for subsequent annotation, this item will exclude multi-mapped ones to be annotated</td></tr><tr><td>Unmapped Reads</td><td>Number of reads that cannot be mapped to the reference genome.</td></tr><tr><td>rRNA Reads</td><td>Number of reads that mapped to the rRNA regions.</td></tr></tbody></table>

### Annotation

Metrics of reads to be annotated by GTF/GFF files.

<table><thead><tr><th width="234">Metric</th><th>Description</th></tr></thead><tbody><tr><td>Exonic</td><td>Number of reads that mapped uniquely to an exonic region and on the same strand of the genome.</td></tr><tr><td>Intronic</td><td>Number of reads that mapped uniquely to an intronic region and on the same strand of the genome.</td></tr><tr><td>Intergenic</td><td>Number of reads that mapped uniquely to an intergenic region and on the same strand of the genome.</td></tr><tr><td>Antisense</td><td>Number of reads mapped to the transcriptome but on the opposite strand of their annotated gene.</td></tr></tbody></table>

### Tissue related

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FG9nsDv6XOFT1M3HSUawE%2Fsummary-tissue-segmentation.PNG?alt=media&amp;token=1d4ad5f1-2182-4755-a15b-20cbbb384fc2" alt=""><figcaption><p>Display and metrics related to tissue-coverage region</p></figcaption></figure>

The tissue segmentation result based on a microscope image is shown on the left, of which the tissue region is covered in purple.

Metrics related to tissue coverage are listed:

<table><thead><tr><th width="244">Metric</th><th>Description</th></tr></thead><tbody><tr><td>Tissue Area</td><td>Tissue area in mm²​.</td></tr><tr><td>Number of MID Under Tissue Coverage</td><td>Number of MID under tissue coverage.</td></tr><tr><td>Fraction MID in Spots Under Tissue</td><td><p>Fraction of MID under tissue over total unique reads.</p><p> (MID Under Tissue / Unique Reads)</p></td></tr></tbody></table>

### Sequencing saturation

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FaflrmI3e43n2ugG3aNeQ%2Fsummary-saturation.PNG?alt=media&amp;token=6b492889-5ca5-4ce8-96ce-5de8a689f67c" alt=""><figcaption><p>Sequencing saturation curves</p></figcaption></figure>

The saturation analysis in the HTML report can assess the overall quality of the sequencing data. In order to improve calculation efficiency, small samples are randomly selected from successfully annotated reads in the bin20 or bin50 dimension. Therefore, the results of multiple runs of the same data may vary slightly. The formulas may not be identical, but the general shape of the curve is consistent.

* Figure 1: As the number of random samples increases, the gene median in the bin20 dimension gradually increases.
* Figure 2: Curves fitted based on Unique Reads data from randomly sampled samples.
* **Figure 3**: Statistics of Unique Reads (reads with unique CID, geneName and MID) in the sampled samples, saturation value = 1-(Unique Reads)/(Total Annotated Reads), as the sampling volume increases, the fitting curve becomes near-flat, indicating that the data tends to be saturated. Whether to add additional tests depends on the overall project design and sample conditions. For example, it is recommended that additional tests be performed on precious samples. The threshold value of 0.8 in the report serves as a reminder for recommended guidance.

{% hint style="info" %}
The x-axis of the three graphs is the same, and the y-axis is divided into saturation value, gene median, and number of Unique Reads.
{% endhint %}

### Information

This item displays the basic information of the input FASTQs,&#x20;

`Species` is from the `--organism` parameter used in `SAW count`, usually referring to the species.

`Tissue` is from the `--tissue` parameter used in `SAW count`.

`Reference` means the reference genome used in `SAW count`, as the same as `Organism`.

`FASTQ` records FASTQ files in `SAW count`, including file prefixes of all input sequencing FASTQs.

## **Square Bin**

This page contains statistics, plots, clustering, UMAP, and differential expression analysis results, at bin dimension. Results come from the analysis based on `<SN>.tissue.gef` file.

### Statistics

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FTBw7queKCFJP2fpyfo2t%2Fsquarebin-table.PNG?alt=media&amp;token=7983e968-17ba-467a-9a05-2ac6d559c52b" alt=""><figcaption><p>Statistics of bins under tissue-coverage region</p></figcaption></figure>

The above table records the statistics for bin20, bin50, and bin200:

<table><thead><tr><th width="247">Item</th><th>Description</th></tr></thead><tbody><tr><td>Bin Size</td><td><p>The size of Bin which is the unit of aggregated DNBs in a squared region.</p><p>i.e. Bin 50 = 50 * 50 DNBs</p></td></tr><tr><td>Mean Reads (per bin)</td><td>Mean number of sequenced reads divided by the number of bins under tissue coverage.</td></tr><tr><td>Median Reads (per bin)</td><td>Median number of sequenced reads divided by the number of bins under tissue coverage (pick the middle value after sorting).</td></tr><tr><td>Mean Gene Type (per bin)</td><td>Mean number of unique gene types divided by the number of bins under tissue coverage.</td></tr><tr><td>Median Gene Type (per bin)</td><td>Median number of unique gene types divided by the number of bins under tissue coverage.</td></tr><tr><td>Mean MID (per bin)</td><td>Mean number of MIDs divided by the number of bins under tissue coverage.</td></tr><tr><td>Median MID (per bin)</td><td>Median number of MIDs divided by the number of bins under tissue coverage</td></tr></tbody></table>

### Plots

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2F0myJ4330BdNAjIrMW5K9%2Fsquarebin-violin.PNG?alt=media&amp;token=3bc6bd15-764c-408a-82f6-9a24d803f873" alt=""><figcaption><p>Distribution plots of MID and gene type</p></figcaption></figure>

Violin plots show the distribution of deduplicated MID count and gene types in each bin.

### **Clustering & UMAP**

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FWELThk8695DM4tDIKYT8%2Fsquarebin-cluster.PNG?alt=media&amp;token=197765f2-711a-44a2-8522-d96a77c0b477" alt=""><figcaption><p>Leiden clustering and UMAP projection</p></figcaption></figure>

Clustering is performed based on `SN.tissue.gef` using the Leiden algorithm. UMAP projections are performed based on `SN.tissue.gef` and colored by automated clustering. The same color is assigned to spots that are within a shorter distance and with similar gene expression profiles.

### Differential expression analysis

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FUEGgoZMtu5XJg1PczJbF%2Fsquarebin-df-table.PNG?alt=media&amp;token=c2d68add-eae9-49a5-a49d-144eea4c79ab" alt=""><figcaption><p>Marker feature table</p></figcaption></figure>

The goal of the differential expression analysis is to identify markers that are more highly expressed in a cluster than the rest of the sample. For each marker, a differential expression test was run between each cluster and the remaining sample. An estimate of the log2 ratio of expression in a cluster to that in other coordinates is Log2 fold-change (L2FC). A value of 1.0 denotes a 2-fold increase in expression within the relevant cluster. Based on a negative binomial test, the p-value indicates the expression difference's statistical significance. The Benjamini-Hochberg method has been used to correct the p-value for multiple testing. Additionally, the top N features by L2FC for each cluster were kept after features in this table were filtered by (Mean UMI counts > 1.0). Grayed-out features have an adjusted p-value >= 0.10 or an L2FC < 0. N (ranges from 1 to 50) is the number of top features displayed per cluster, which is set to limit the amount of table entries displayed to 10,000. N=%10,000/K^2 where K is the number of clusters. Click on a column to sort by that value, or search a gene of interest.

{% hint style="info" %}
When the values of L2FC in the marker feature table are blank, "infinity" and "-infinity", the analysis results are normal. These conditions are well explained below.
{% endhint %}

The calculation of L2FC is related to the expression number of cells of a certain gene in the case group and the control group. Since the calculation of L2FC uses the natural logarithm as the base, when the expression relationship has extremely high or low values, the three special values, none, "inf" and "-inf", will appear. The screenshot below uses inf and a constant to make a simple demonstration.

<figure><img src="https://content.gitbook.com/content/33hMinoADCychkEZKBYW/blobs/JF6tHoVLwAFozloPxQG0/image.png" alt="" width="563"><figcaption><p>An example in Notebook using Python</p></figcaption></figure>

{% hint style="warning" %}
The p-values should be increasing as the list descends (with a maximum of 1), infinitely close to 0.&#x20;

If you find that the p-value is 0 in the result table, it may be because the calculated differential expression feature is extremely significant, leading to an extremely small p-value. This can exceed the limit of the data type (usually `float64`, depending on the basic computing package), resulting in a situation that cannot be expressed in scientific notation.
{% endhint %}

## **Cell Bin**

This page contains results of statistics, plots, clustering, UMAP, and differential expression analysis, at cellbin dimension. Cell border expanding is automatically performed during `SAW count` and `SAW realign`, which means the contents of "Cell Bin" tab are based on `SN.adjusted.cellbin.gef`.&#x20;

{% hint style="warning" %}
When it comes to `--adjusted-distance=0` in `SAW realign`, all contents of this tab are based on `SN.cellbin.gef`.&#x20;
{% endhint %}

### Statistics

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2F6ri5WSxmRSO3B2h5Azqn%2Fcellbin-table.PNG?alt=media&amp;token=59c9b600-8bf8-41b9-bf87-cac1674fad63" alt=""><figcaption><p>Detailed statistics of cellbin</p></figcaption></figure>

The above table records the statistics of cellbin:

<table><thead><tr><th width="295">Item</th><th>Description</th></tr></thead><tbody><tr><td>Cell Count</td><td>Number of cells.</td></tr><tr><td>Mean Cell Area</td><td>Mean cell area, in pixes.</td></tr><tr><td>Median Cell Area</td><td>Median cell area, in pixes.</td></tr><tr><td>Mean Gene Type</td><td>Mean gene types per cell.</td></tr><tr><td>Median Gene Type</td><td>Median gene types per cell.</td></tr><tr><td>Mean MID</td><td>Mean MID count per cell.</td></tr><tr><td>Median MID</td><td>Median MID count per cell.</td></tr></tbody></table>

### Plots

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2F5GFtrl1Wsd9UJBM7o7BP%2Fcellbin-violin.PNG?alt=media&amp;token=693588a2-115e-4cad-be73-bd60b85e365b" alt=""><figcaption><p>Distribution plots of MID and gene type</p></figcaption></figure>

Violin plots show the distribution of deduplicated MID count, gene types and cell area in the cellbin.

### **Clustering & UMAP**

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FGIuC9mG5eJmLenIIiJQw%2Fcellbin-cluster.PNG?alt=media&amp;token=900c6504-26c9-455d-980c-13aed673d074" alt=""><figcaption><p>Leiden clustering and UMAP projection</p></figcaption></figure>

Clustering is performed based on `SN.adjusted.cellbin.gef` or `SN.cellbin.gef`, using the Leiden algorithm. UMAP projections are performed based on `SN.adjusted.cellbin.gef` or `SN.cellbin.gef`, and colored by automated clustering. The same color is assigned to spots that are within a shorter distance and with similar gene expression profiles.

### Differential expression analysis

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FH2FPV5kezfy6F9wLNMSv%2Fcellbin-df-table.PNG?alt=media&amp;token=ba92880f-cbdc-4ba1-8ed9-dfbf03e7ec20" alt=""><figcaption><p>Marker feature table</p></figcaption></figure>

The goal of the differential expression analysis is to identify markers that are more highly expressed in a cluster than the rest of the sample. For each marker, a differential expression test was run between each cluster and the remaining sample. An estimate of the log2 ratio of expression in a cluster to that in other coordinates is Log2 fold-change (L2FC). A value of 1.0 denotes a 2-fold increase in expression within the relevant cluster. Based on a negative binomial test, the p-value indicates the expression difference's statistical significance. The Benjamini-Hochberg method has been used to correct the p-value for multiple testing. Additionally, the top N features by L2FC for each cluster were kept after features in this table were filtered by (Mean UMI counts > 1.0). Grayed-out features have an adjusted p-value >= 0.10 or an L2FC < 0. N (ranges from 1 to 50) is the number of top features displayed per cluster, which is set to limit the amount of table entries displayed to 10,000. N=%10,000/K^2 where K is the number of clusters. Click on a column to sort by that value, or search a gene of interest.

{% hint style="info" %}
Interpretation for exceptional cases related to differential expression analysis can be found under [Square Bin](#differential-expression-analysis) part.
{% endhint %}

## **Image**

### Image information

Basic information about the microscopic staining image, usually involving microscope settings.

### QC

<table><thead><tr><th width="232">Metric</th><th>Description</th></tr></thead><tbody><tr><td>Image QC version</td><td>The version of image QC module.</td></tr><tr><td>QC Pass</td><td>Whether the image(s) passed image QC quality check.</td></tr><tr><td>Trackline Score</td><td>Reference score for trackline detection.</td></tr></tbody></table>

### Stitching

<table><thead><tr><th width="251">Metric</th><th>Description</th></tr></thead><tbody><tr><td>Template Source Row No.</td><td>The row number of the template FOV used for predicting the entire template.</td></tr><tr><td>Template Source Column No.</td><td>The column number of the template FOV used for predicting the entire template.</td></tr><tr><td>Global Height</td><td>Height of the stitched image.</td></tr><tr><td>Global Width</td><td>Width of the stitched image.</td></tr></tbody></table>

### Registration

<table><thead><tr><th width="256">Metric</th><th>Description</th></tr></thead><tbody><tr><td>ScaleX</td><td>The lateral scaling between image and template.</td></tr><tr><td>ScaleY</td><td>The longitudinal scaling between image and template.</td></tr><tr><td>Rotation</td><td>The rotation angle of the image relative to the template.</td></tr><tr><td>Flip</td><td>Whether the image is flipped horizontally.</td></tr><tr><td>Image X Offset</td><td>Offset between image and matrix in x direction.</td></tr><tr><td>Image Y Offset</td><td>Offset between image and matrix in y direction</td></tr><tr><td>Counter Clockwise Rotation</td><td>Counter clockwise rotation angle.</td></tr><tr><td>Manual ScaleX</td><td>The lateral scaling based on image center (manual-registration).</td></tr><tr><td>Manual ScaleY</td><td>The longitudinal scaling based on image center (manual-registration).</td></tr><tr><td>Manual Rotation</td><td>The rotation angle based on image center (manual-registration).</td></tr><tr><td>Matrix X Start</td><td>Gene expression matrix offset in x direction by DNB numbers.</td></tr><tr><td>Matrix Y Start</td><td>Gene expression matrix offset in y direction by DNB numbers.</td></tr><tr><td>Matrix Height</td><td>Gene expression matrix height.</td></tr><tr><td>Matrix Width</td><td>Gene expression matrix width.</td></tr></tbody></table>

## **Microorganism**

{% hint style="warning" %}
Here is an another FFPE tissue sample of mouse lung which is especially for microorganism analysis.
{% endhint %}

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FR2ieToIeym3xmZPg7M5y%2Fmicroorganism-expression.PNG?alt=media&amp;token=00739503-e147-4c44-bc8e-b07416829bc2" alt=""><figcaption><p>Microorganism heatmap under tissue region and four key metrics</p></figcaption></figure>

The distribution plot of microorganism spatial expression, on the left, shows MID count at bin20.

### Denoising

<table><thead><tr><th width="244">Metric</th><th>Description</th></tr></thead><tbody><tr><td>Total Reads</td><td>Total number of input reads.</td></tr><tr><td>Non-Host Source Reads</td><td>Number of reads that can not be aligned to the host genome.</td></tr><tr><td>Host Source Reads</td><td>Number of reads that can be aligned to the host genome during denoising.</td></tr></tbody></table>

### Taxonomic Classification

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2F9eDUA22N4mcZEcxs1ymy%2Fmicroorganism-taxonnomy.PNG?alt=media&amp;token=1272234b-576a-4ffd-94c1-9835e8d907ee" alt=""><figcaption><p>Mapping results of Bowtie2 and Kraken2</p></figcaption></figure>

<table><thead><tr><th width="264">Metric</th><th>Description</th></tr></thead><tbody><tr><td>Non-Host Source Reads</td><td>Number of reads that can not be aligned to the host genome.</td></tr><tr><td>Bacteria, Fungi and Viruses MIDs</td><td>Number of unique mRNA molecular assigned to bacteria, fungi or viruses.</td></tr><tr><td>Bacteria, Fungi and Viruses Duplication</td><td>Number of assigned reads that have been corrected due to duplicated MID.</td></tr><tr><td>Other Microbes or Host-Suspicious</td><td>Number of reads assigned to other microbes (exclude bacteria, fungi and viruses) or host.</td></tr><tr><td>Unclassified Reads</td><td>Number of unclassified reads.</td></tr></tbody></table>

### Microbes Proportion (Phylum)

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2Fq3FbN7B7soAe0zkhr7p2%2Fmicroorganism-proportion.PNG?alt=media&amp;token=7bee6a62-73de-46b9-a93d-18100bdb00f7" alt=""><figcaption><p>Microbes proportion at phylum level</p></figcaption></figure>

The main proportion of microbes at the phylum level.

*\*the same for other classifications*

## Summary-Protein

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FVSKBDG5w8yCqIfCX2lvI%2Fsummary-protein-expression.PNG?alt=media&amp;token=2c3d4cfa-d52e-44bc-a002-85bd0a7b07b0" alt=""><figcaption><p>Expression heatmap and four key metrics</p></figcaption></figure>

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2F4BBUUHqngH2MLn2L4B6z%2Fsummary-protein-image.PNG?alt=media&amp;token=7d531a0f-c9ea-4ebc-b215-d360e1e0bdba" alt=""><figcaption><p>Display of microscope image</p></figcaption></figure>

The spatial protein expression distribution plot, containing all bins and bins under tissue, on the left, shows MID count at each bin20.

`Total Reads` is the total sequencing reads of input sequencing ADT FASTQs. `Valid CID reads` represents the number of reads with CIDs matching the mask file, with MIDs passing QC. `Valid PID reads` represents the number of reads that are mapped to the PID sequence in the protein panel. `Unique PID reads` represents the total number of unique protein reads (PID reads whose MIDs are different).

### Key metrics

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FlzeVKm4Y2j14FfBwQtEP%2Fsummary-protein-key-metrics.PNG?alt=media&amp;token=f35d5826-068a-4fcb-ae56-30f1446f7337" alt=""><figcaption><p>Details and sunburst plot of key metrics</p></figcaption></figure>

Key metrics of the data are listed:

<table><thead><tr><th width="234">Metrics</th><th>Description</th></tr></thead><tbody><tr><td>Total Reads</td><td>Total number of sequenced reads</td></tr><tr><td>Valid CID Reads</td><td>Number of reads with CIDs matching the mask file and with MIDs passing QC</td></tr><tr><td>Invalid CID Reads</td><td>Number of reads with CIDs that cannot be matched with the mask file</td></tr><tr><td>Valid PID Reads</td><td>Valid CID reads that mapped to the protein sequence in the protein sequence database (protein panel)</td></tr><tr><td>Invalid PID Reads</td><td>Valid CID reads that can not be mapped to the protein sequence in the protein sequence database (protein panel)</td></tr><tr><td>Unique PID Reads</td><td>Total number of unique protein reads (PID reads whose MIDs are different)</td></tr><tr><td>Sequencing Saturation</td><td>Number of PID reads with duplicated MID</td></tr></tbody></table>

### Sequencing saturation

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FvKK2yAZtIDPuC3A66Ug5%2Fsummary-protein-saturation.PNG?alt=media&amp;token=2d36e4ed-8fd5-41c8-a5ef-c87fd8fda99d" alt=""><figcaption></figcaption></figure>

The saturation analysis in the HTML report can assess the overall quality of the sequencing data. In order to improve calculation efficiency, small samples are randomly selected from successfully annotated reads in the bin20 or bin50 dimension. Therefore, the results of multiple runs of the same data may vary slightly. The formulas may not be identical, but the general shape of the curve is consistent.

* Figure 1: Curves fitted based on Unique Reads data from randomly sampled samples.
* **Figure 2**: Statistics of Unique Reads (reads with unique CID, PID and MID) in the sampled samples, saturation value = 1-(Unique Reads)/(Valid PID Reads), as the sampling volume increases, the fitting curve becomes near-flat, indicating that the data tends to be saturated. Whether to add additional tests depends on the overall project design and sample conditions. For example, it is recommended that additional tests be performed on precious samples. The threshold value of 0.8 in the report serves as a reminder for recommended guidance.

### Protein correlations

Spearman correlation (in bin20 or bin50) between raw antibody counts under tissue, except isotype. Antibodies are clustered based on Spearman correlation coefficient.

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FyG7iDpIFGSzEf9CNurVA%2Fsummary-protein-correlation.PNG?alt=media&amp;token=bef6be48-8e14-430b-a4fe-c741b1c8970b" alt=""><figcaption><p>Correlation plot within proteins</p></figcaption></figure>

### Gene : protein correlations

Spearman correlation (in bin20 or bin50) between raw gene counts and raw antibody counts under tissue, where antibody has at least one marker gene in the protein panel.

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FEBKK5b4HfVj1nhKqTqY2%2Fsummary-protein-gene-correlation.PNG?alt=media&amp;token=07e42af6-e396-44af-984b-d2bbb55abc6a" alt=""><figcaption></figcaption></figure>

### Histogram of portein counts

Distribution of spot numbers vs log-scaled MID count (in bin20 or bin50).

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2F5lLS8gu1gPlVX75KLjZT%2Fsummary-protein-histogram.PNG?alt=media&amp;token=73433e38-9cda-4321-9d00-71150fa0b555" alt=""><figcaption><p>Histogram plot between spot numbers and log-scaled MID count</p></figcaption></figure>

### Information

This item displays the basic information of the input FASTQs,&#x20;

`Species` is from the `--organism` parameter used in `SAW count`, usually referring to the species.

`Tissue` is from the `--tissue` parameter used in `SAW count`.

`Reference` means the reference genome used in `SAW count`, as the same as `Organism`.

`FASTQ` records  FASTQ files in `SAW count`, including file prefixes of all input sequencing ADT FASTQs.

## Square Bin-Protein

This page contains results of statistics, plots, clustering, UMAP, and differential expression analysis, at bin dimension. Results come from the analysis based on `<SN>.protein.tissue.gef` file.

### Statistics

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2F13QKGhLp3KNFi99N8O3x%2Fsquarebin-protein-table.PNG?alt=media&amp;token=1b20e459-5b27-4935-bc5a-4398ad118980" alt=""><figcaption><p>Statistics of bins under tissue-coverage region</p></figcaption></figure>

<table><thead><tr><th width="247">Item</th><th>Description</th></tr></thead><tbody><tr><td>Bin Size</td><td><p>The size of Bin which is the unit of aggregated DNBs in a squared region.</p><p>i.e. Bin 50 = 50 * 50 DNBs</p></td></tr><tr><td>Mean Reads (per bin)</td><td>Mean number of sequenced reads divided by the number of bins under tissue coverage.</td></tr><tr><td>Median Reads (per bin)</td><td>Median number of sequenced reads divided by the number of bins under tissue coverage (pick the middle value after sorting).</td></tr><tr><td>Mean MID (per bin)</td><td>Mean number of MIDs divided by the number of bins under tissue coverage.</td></tr><tr><td>Median MID (per bin)</td><td>Median number of MIDs divided by the number of bins under tissue coverage</td></tr></tbody></table>

### Plots

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2Fnm9t6csUdnTxl1hhvSFw%2Fsquarebin-protein-violin.PNG?alt=media&amp;token=cba78f86-848c-4743-8867-10dd23106a11" alt=""><figcaption><p>Distribution plots of MID</p></figcaption></figure>

Violin plots show the distribution of deduplicated MID count in each bin.

### Clustering & UMAP

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FnluLe66si4S3mDLBnmvr%2Fbin20-protein-cluster.png?alt=media&amp;token=a7f9ed17-0fa7-46ba-af5e-811f0d1e13f4" alt=""><figcaption><p>Leiden clustering and UMAP projection</p></figcaption></figure>

Clustering is performed based on `SN.protein.tissue.gef` using the Leiden algorithm. UMAP projections are performed based on `SN.protein.tissue.gef` and colored by automated clustering. The same color is assigned to spots that are within a shorter distance and with similar gene expression profiles.

## Cell Bin-Protein

This page contains results of statistics, plots, clustering, UMAP, and differential expression analysis, at cellbin dimension. Cell border expanding is automatically performed during `SAW count` and `SAW realign`, which means the contents of "Cell Bin" tab are based on `SN.protein.adjusted.cellbin.gef`.&#x20;

{% hint style="warning" %}
When it comes to `--adjusted-distance=0` in `SAW realign`, all contents of this tab are based on `SN.protein.cellbin.gef`.&#x20;
{% endhint %}

### Statistics

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FElY8AZn35D1fj2P21rSa%2Fcellbin-protein-table.PNG?alt=media&amp;token=b77dd86d-6293-490f-adf7-c2eeb26dcd3b" alt=""><figcaption></figcaption></figure>

The above table records the statistics of cellbin:

<table><thead><tr><th width="247">Item</th><th>Description</th></tr></thead><tbody><tr><td>Cell Count</td><td>Number of cells.</td></tr><tr><td>Mean Cell Area</td><td>Mean cell area, in pixes.</td></tr><tr><td>Median Cell Area</td><td>Median cell area, in pixes.</td></tr><tr><td>Mean MID</td><td>Mean MID count per cell.</td></tr><tr><td>Median MID</td><td>Median MID count per cell.</td></tr></tbody></table>

### Plots

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FtsEwI2zJrLzuZQ2pUamA%2Fcellbin-protein-violin.PNG?alt=media&amp;token=abd40e10-8dba-40b1-981f-420a9c5da8d4" alt=""><figcaption></figcaption></figure>

Violin plots show the distribution of deduplicated MID count and cell area in the cellbin.

### Clustering & UMAP

<figure><img src="https://1692821827-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F33hMinoADCychkEZKBYW%2Fuploads%2FhcFsIK49402MobpiaxIz%2Fcellbin-protein-cluster.png?alt=media&amp;token=57b98f86-d01c-4e7c-a502-3cbd4e23cd2e" alt=""><figcaption><p>Leiden clustering and UMAP projection</p></figcaption></figure>

Clustering is performed based on `SN.protein.adjusted.cellbin.gef` or `SN.protein.cellbin.gef`, using the Leiden algorithm. UMAP projections are performed based on `SN.protein.adjusted.cellbin.gef` or `SN.protein.cellbin.gef`, and colored by automated clustering. The same color is assigned to spots that are within a shorter distance and with similar gene expression profiles.

## Analysis (no longer available since 8.1.2)

This page contains results of clustering from proteome & transcriptome joint analysis, marker selection, and correlation between genes and proteins.&#x20;

### Multiomics clustering & UMAP

Clustering and UMAP projections are performed based on the latent space generated by totalVI.

### Top markers by cluster

Heatmaps of top <=3 gene and protein markers per Leiden cluster from gene-protein jointly analysis. These features are filtered after one-vs-all differential expression analysis, following these rules:

* For gene, Bayes factor > 1 and  expression proportion greater than 10% in the cluster;
* For protein, Bayes factor > 0.7.

## **Alerts**

Thresholds are set for several important statistical indicators. If the analysis results are abnormal, an alert message will be displayed at the top of the HTML report.

{% hint style="warning" %}
Here is an abnormal exmple data just for display.
{% endhint %}

<figure><img src="https://content.gitbook.com/content/33hMinoADCychkEZKBYW/blobs/2Lz1FuBjt4zttV3346ZD/image.png" alt=""><figcaption><p>Alert information</p></figcaption></figure>
