Difference between revisions of "Data Monitoring Procedures"

Latest revision as of 13:18, 2 November 2023

Master List of File / Database / Webpage Locations

Run Conditions

Online Run-by-run condition files (B-field, current, etc.): /work/halld/online_monitoring/conditions/
Offline monitoring run conditions (software versions, jana config): /group/halld/data_monitoring/run_conditions/
Run Info vers. 1
Run Info vers. 2
RCDB

Monitoring Output Files

Run Periods 201Y-MM is for example 2015-03, launch ver verVV is for example ver15
Online monitoring histograms: /work/halld/online_monitoring/root/
Offline monitoring histogram ROOT files (merged): /work/halld/data_monitoring/RunPeriod-201Y-MM/verVV/rootfiles
individual files for each job (ROOT, REST, log, etc.): /volatile/halld/offline_monitoring/RunPeriod-201Y-MM/verVV/

Monitoring Database

Accessing monitoring database (on ifarm): mysql -u datmon -h hallddb.jlab.org data_monitoring

Monitoring Webpages

SciComp Job Links

Main

Documentation

Job Tracking

Procedures: Overview

Online Monitoring: During Experimental Running

After every run is finished, a ROOT file containing histograms from the online monitoring system and a file containing some run conditions are copied to directories under /work/halld/online_monitoring . A cronjob running in the counting house performs this function.

This ROOT file is processed similarly to the offline monitoring results, and are made available under the same webpages as "ver00" of the relevant run period.

For more details on the online monitoring system, see this page.

Offline Monitoring and Reconstruction: During Experimental Running

During experimental running, the following offline monitoring procedures should be performed, each with a different gxprojN account, so that they don't interfere with each other:

Incoming: Monitor the first 5 files of each newly-recorded run as soon as it hits the tape.
Monitoring Launches: Every two weeks, do a monitoring launch over the first 5 files of all runs currently available on the tape.
Initial Reconstruction Launch: As soon as a new group (e.g. ~100 runs) of data is initially semi-well calibrated, do a preliminary full reconstruction launch over all files in that group.
- We can add user analysis plugins to this launch, including those with ROOT TTree output, provided that they work and don't take much memory.

Note that the monitoring is limited to the first 5 files of each run, because data is being recorded to tape at a faster rate than the monitoring can keep up with. Also, during the experimental run, each run will only be fully-reconstructed once, because it will be difficult enough to keep up with the incoming data.

Offline Monitoring and Reconstruction: After Experimental Running

After experimental running, the following offline monitoring procedures should be performed, each with a different gxprojN account, so that they don't interfere with each other:

Monitoring Launches: Every two weeks, do a monitoring launch over the first 5 files of all runs currently available on the tape.
Initial Reconstruction Launch: As soon as a new group (e.g. ~100 runs) of data is initially semi-well calibrated, do a preliminary full-reconstruction launch over all files in that group.
- We can add user analysis plugins to this launch, including those with ROOT TTree output, provided that they work and don't take much memory.
Further Reconstruction Launches: Every ~3 months, if there have been significant improvements to the reconstruction / calibrations, do a new full-reconstruction launch over all of the data.
- We can add user analysis plugins to this launch, including those with ROOT TTree output, provided that they work and don't take much memory.

Note that the monitoring is limited to the first 5 files of each run, since there will be a significant amount of data.

Saving to Tape (Write-through Cache): Monitoring Launches

All job output will be directly written to the write-thru cache. However, only the following will be saved to tape:

REST files: All files.
ROOT files: One merged file per run.
- After merge, the individual files are deleted (so they won't be saved).
Job stdout/stderr: One tarball per run
- After launch analysis, the log files are deleted (so they won't be saved).
Browser png's: One tarball per launch

Saving to Tape (Write-through Cache): Full Reconstruction Launches

REST files: All files.
ROOT files: All files, AND one merged file per run.
Job stdout/stderr: One tarball per run
- After launch analysis, a tarball is created and the individual log files are deleted (so they won't be saved).
Browser png's: One tarball per launch

Procedures: Details

On- and Offline Monitoring Data Validation

Software Tests

Software Test: Experimental Data Reconstruction
- Test Results

@@ Line 19: / Line 19: @@
 === Monitoring Webpages ===
-*[https://halldweb.jlab.org/cgi-bin/data_monitoring/monitoring/plotBrowser.py Plot Browser]
+*[https://halldweb.jlab.org/wiki/index.php/Monitoring_webpage_help Help]
+*[https://halldweb.jlab.org/data_monitoring/Plot_Browser.html Plot Browser]
 *[https://halldweb.jlab.org/cgi-bin/data_monitoring/monitoring/runBrowser.py Run Browser]
 *[https://halldweb.jlab.org/cgi-bin/data_monitoring/monitoring/versionBrowser.py Version Browser]
 *[https://halldweb.jlab.org/cgi-bin/data_monitoring/monitoring/timeSeries.py Time Series]
 *[https://halldweb.jlab.org/data_monitoring/launch_analysis/ Launch Analysis]
+*[https://halldweb.jlab.org/cgi-bin/data_monitoring/monitoring/recontestBrowser.py Recon Tests]
 == SciComp Job Links ==
@@ Line 29: / Line 31: @@
 * [https://scicomp.jlab.org/scicomp/ Scientific Computing Home Page]
 * [https://scicomp.jlab.org/scicomp/#/auger/jobs Auger Job Status Page]
-* [https://scicomp.jlab.org/scicomp/#/jasmine/jobs Jasmine Tape Job Status Page]
+* [https://scicomp.jlab.org/scicomp/#/jasmine/jobs JasMine Tape Job Status Page]
 === Documentation ===
 * [https://scicomp.jlab.org/docs/batch Batch System]
 * [https://scicomp.jlab.org/docs/storage Mass Storage System]
+* [https://scicomp.jlab.org/docs/write-through-cache Write-Through Cache]
 * [https://scicomp.jlab.org/docs/swif SWIF]
 * [https://scicomp.jlab.org/docs/swif-cli SWIF Command Line]
@@ Line 45: / Line 48: @@
 == Procedures: Overview ==
+=== Online Monitoring: During Experimental Running ===
+After every run is finished, a ROOT file containing histograms from the online monitoring system and a file containing some run conditions are copied to directories under /work/halld/online_monitoring . A cronjob running in the counting house performs this function.
+This ROOT file is processed similarly to the offline monitoring results, and are made available under the same webpages as "ver00" of the relevant run period.
+For more details on the online monitoring system, see [https://halldweb.jlab.org/hdops/wiki/index.php/Online_Monitoring_Shift  this page].
 === Offline Monitoring and Reconstruction: During Experimental Running ===
@@ Line 50: / Line 61: @@
 During experimental running, the following offline monitoring procedures should be performed, each with a different gxprojN account, so that they don't interfere with each other:
-# Monitor the first <span style="color:red">20</span> files of each newly-recorded run as soon as it hits the tape.
+# '''Incoming:''' Monitor the first <span style="color:red">5</span> files of each newly-recorded run as soon as it hits the tape.
-# Every <span style="color:red">two</span> weeks, do a monitoring launch over the first <span style="color:red">20</span> files of all runs currently available on the tape.
+# '''Monitoring Launches:''' Every <span style="color:red">two</span> weeks, do a monitoring launch over the first <span style="color:red">5</span> files of all runs currently available on the tape.
-# As soon as a new group (e.g. <span style="color:red">~100</span> runs) of data is initially semi-well calibrated, do a preliminary full reconstruction launch over all files in that group.
+# '''Initial Reconstruction Launch:''' As soon as a new group (e.g. <span style="color:red">~100</span> runs) of data is initially semi-well calibrated, do a preliminary full reconstruction launch over all files in that group.
 #* We can add user analysis plugins to this launch, including those with ROOT TTree output, provided that they work and don't take much memory.
-Note that the monitoring is limited to the first <span style="color:red">20</span> files of each run, because data is being recorded to tape at a faster rate than the monitoring can keep up with.  Also, during the experimental run, each run will only be fully-reconstructed once, because it will be difficult enough to keep up with the incoming data.
+Note that the monitoring is limited to the first <span style="color:red">5</span> files of each run, because data is being recorded to tape at a faster rate than the monitoring can keep up with.  Also, during the experimental run, each run will only be fully-reconstructed once, because it will be difficult enough to keep up with the incoming data.
 === Offline Monitoring and Reconstruction: After Experimental Running ===
@@ Line 61: / Line 72: @@
 After experimental running, the following offline monitoring procedures should be performed, each with a different gxprojN account, so that they don't interfere with each other:
-# Every two weeks, do a monitoring launch over the first 20 files of all runs currently available on the tape.
+# '''Monitoring Launches:''' Every two weeks, do a monitoring launch over the first <span style="color:red">5</span> files of all runs currently available on the tape.
-# As soon as a new group (e.g. ~100 runs) of data is initially semi-well calibrated, do a preliminary full-reconstruction launch over all files in that group.
+# '''Initial Reconstruction Launch:''' As soon as a new group (e.g. <span style="color:red">~100</span> runs) of data is initially semi-well calibrated, do a preliminary full-reconstruction launch over all files in that group.
-# Every three months, if there have been significant improvements to the reconstruction / calibrations, do a new full-reconstruction launch over all of the data.
+#* We can add user analysis plugins to this launch, including those with ROOT TTree output, provided that they work and don't take much memory.
+# '''Further Reconstruction Launches:''' Every <span style="color:red">~3</span> months, if there have been significant improvements to the reconstruction / calibrations, do a new full-reconstruction launch over all of the data.
 #* We can add user analysis plugins to this launch, including those with ROOT TTree output, provided that they work and don't take much memory.
-Note that the monitoring is limited to the first 20 files of each run, since there will be a significant amount of data.
+Note that the monitoring is limited to the first <span style="color:red">5</span> files of each run, since there will be a significant amount of data.
-=== Saving to Tape (Write-thru Cache): Monitoring Launches ===
+=== Saving to Tape (Write-through Cache): Monitoring Launches ===
+All job output will be directly written to the write-thru cache. However, only the following will be saved to tape:
 * REST files: All files.
 * ROOT files: One merged file per run.
-* Job stdout/stderr: None
+** After merge, the individual files are deleted (so they won't be saved).
+* Job stdout/stderr: One tarball per run
+** After launch analysis, the log files are deleted (so they won't be saved).
 * Browser png's: One tarball per launch
-=== Saving to Tape (Write-thru Cache): Full Reconstruction Launches ===
+=== Saving to Tape (Write-through Cache): Full Reconstruction Launches ===
 * REST files: All files.
-* ROOT files: All files, AND one merged file per run.
+* ROOT files: All files, <span style="color:blue">AND</span> one merged file per run.
 * Job stdout/stderr: One tarball per run
+** After launch analysis, a tarball is created and the individual log files are deleted (so they won't be saved).
 * Browser png's: One tarball per launch
 == Procedures: Details ==
@@ Line 85: / Line 102: @@
 * [[Offline_Monitoring_Archived_Data | Offline Monitoring: Running Over Archived Data]]
 * [[Offline_Monitoring_Post_Processing | Offline Monitoring: Post-Processing]]
+* [[DEPRECATED_Offline_Monitoring_Archived_Data | DEPRECATED (Except plots): Offline Monitoring: Running Over Archived Data]]
+* [[DSelector_SWIF_Jobs | DSelector SWIF Jobs]]
+* [[Merging_Analysis_Trees | Analysis Launch: Merging Trees]]
+=== On- and Offline Monitoring Data Validation===
+* [[Offline_Monitoring_Data_Validation | Offline Monitoring: Data Validation]]
+* [[Offline_Monitoring_Data_Validation_PrimEx | Offline Monitoring: Data Validation of PrimEx data]]
+* [[Offline_Monitoring_Data_Validation_CPP | Offline Monitoring: Data Validation of CPP data]]
+* [[Online_Monitoring_Data_Validation | Online Monitoring: Data Validation]]
+== Software Tests ==
+* [[Software_Test_Data_Recon | Software Test: Experimental Data Reconstruction]]
+** [https://halldweb.jlab.org/recon_test/ Test Results]