The PyRETIS analysis application

The PyRETIS analysis application, pyretis analyse, is used to analyse the results from PyRETIS simulations.

See also

Simulation outputs describes the files a run writes and how to read the report this application produces – including how to tell a converged simulation from one that has not sampled enough. Calculating the rate explains how the rate constant itself is put together, and how the routes differ between TIS, RETIS, infinite swapping and the (RE)PPTIS family.

The general syntax for executing is:

pyretis analyse [-h] [-i INPUT] [-V]

where INPUT is the input file for the analysis. This file is similar to the input file to the PyRETIS application with some differences:

  1. It contains the actual number of cycles/steps completed, specified by the endcycle keyword.

  2. It contains the directory from where the simulation was executed, specified by the exe_path keyword.

  3. It contains the number of particles in the simulation, specified by the npart keyword.

  4. It contains settings for the analysis, defined via the analysis section.

Note

The analysis reads the SAME input file the simulation ran from (pyretis analyse -i retis.toml), or the run’s own output.toml (pyretis analyse -i output.toml), which records that input with every default resolved: the analysis takes the input sections from it and the run directory it is run in (see Calculating the rate).

Input arguments

Table 53 Description of input arguments for pyretis analyse

Argument

Description

-h, –help

Show the help message and exit.

-i INPUT, –input INPUT

Location of the input file containing analysis settings.

-V, –version

Show the version number and exit.

-skipb N, –skip-begin N

Discard the first N cycles (equilibration). Optional; without it nothing is discarded.

-skipe N, –skip-end N

Discard the last N cycles, for instance after an interrupted run. Optional; without it nothing is discarded.

Choosing which cycles to analyse

By default every sampled cycle is analysed. Two optional arguments narrow that window:

pyretis analyse -i input.toml -skipb 500
pyretis analyse -i input.toml -skipb 500 -skipe 100

The first discards the first 500 cycles – the usual reason being that the simulation had not yet equilibrated. The second additionally discards the last 100, which is what you want when a run was cut short (a wall-time kill mid-cycle) or when you are comparing two runs over the same window.

Both are counted in cycles, they can be used together or on their own, and they override skip_initial_cycles and skip_final_cycles in the input file, so you can try several windows without editing anything. The analysis log states which cycles it dropped.

The chosen window applies to everything the analysis reports: the crossing probabilities, the rate, and the acceptance ratios below.

Independently of the window, each ensemble’s statistics start at its first accepted, non-initialisation path. The paths an ensemble holds before that are the ones initiation put there (ld, ki or is) and they are not samples of the ensemble, so counting them would bias the crossing probability towards whatever the initial path happened to do. An ensemble that has not yet been sampled at all is reported as having no accepted paths to analyse rather than given a number built from its initial path – if you see that, the run was too short to reach that ensemble.

The same holds for a path marked hc: an ensemble held it at a restart that changed [tis] high_accept, so it was accepted under the other acceptance rule. Its rows from the restart on are left out, and the ensemble counts again from its first path accepted after the restart; its rows from before the restart stay in.

The WHAM analysis, and every tool that builds the combined path-data matrix from the per-ensemble output (the training set and the interface optimiser read the same matrix), leaves out the rows of the initialisation paths and of the paths marked hc by the same rule. Its skip counts paths, the initialisation paths included: -skipb 500 discards the first 500 paths of the run by path number, among the paths the ensembles do not hold at the end of the run, and the rows of the initialisation paths among the paths that remain are left out. A run whose initialisation paths all fall within the skipped paths gives the matrix of every row after that skip.

A literal infswap_data.txt or infretis_data.txt in the run directory is read in place of the per-ensemble output. Its rows record no move, so every row of it is kept, the rows of the initialisation paths included; its skip discards its first rows. The merge of several runs (pyretis.analysis.combine_data) writes <out>.txt, combo.txt by default, in this layout, and its interfaces to <out>.toml. Copied into a directory as infswap_data.txt, with <out>.toml as output.toml, the merge is read in this way. The merge reads the paths of each run as the analysis does, and its skip counts them in the same way.

Stone skipping without high acceptance ([tis] high_accept = false) samples each path with weight 1. When the rows of such an ensemble carry the crossing count of each path as HA-weight, and the counts differ from row to row, every weighted average over them is biased, so the analysis leaves those cycles out and its log names them. The run’s output.toml records, as [current] ss_weight_from_cycle, the first cycle whose rows carry the weight of the rule each path was sampled with, and, under [current] ss_weight_before, the move list and layout the earlier rows were written with and the paths held at that cycle. The rows before it are checked, and so is every row of a run whose output.toml does not record it.

Where no weight decides where a path goes (tis, pptis, repptis), only an affected ensemble loses those cycles. In a retis run the weights also decided, through the swaps, where each path went, so every ensemble loses them, up to the last cycle in which a path held at the boundary is still held. The cycles after that start from the paths the earlier weights left behind, as a run starts from its initial paths; give them an equilibration skip (-skipb) of their own if they need one. An ensemble with no cycle left is refused with an error; continuing the run with restart adds cycles that can be analysed. The WHAM analysis, and every tool that reads the combined path-data matrix, sums every cycle of a path, so it refuses such a run.

Two cases leave no trace in the run’s files: a change of [tis] high_accept that output.toml does not record (a state file edited by hand and resumed from), and rows appended after the boundary by a version of PyRETIS that writes the crossing count.

Acceptance ratios

The report’s Pathensemble data table gives, per ensemble, the fraction of attempted moves that were accepted – separately for the shooting-like moves and for the swapping moves.

These are counted from moves.txt – one line per attempted move: the cycle, the move that ran, and the status it ended in (ACC, or a rejection code such as BWI or NCR; see Path types and rejection reasons). It is the only record of a rejection. pathensemble.txt beside it says which path each ensemble holds, so every one of its lines is an accepted path by construction – counting those would report every move as accepted. See moves.txt – what the sampler attempted for the file’s columns.

A cycle in which two ensembles were picked together is a swap, whatever shooting move those ensembles are configured with, and is recorded in both partners: under the + code in the lower-index one and the - code in the higher. The pair of codes says which swap it was – s+/s- for the zero swap between [0^-] and [0^+], p+/p- for the PPTIS/REPPTIS swap between adjacent ensembles. That is what makes the swap ratio computable at all.

Two further columns report what counting acceptances cannot.

HA eff. is for the high-acceptance moves – wire fencing and stone skipping. Those are built so that a trial is accepted almost every time; what varies is the weight the accepted path carries, its crossing count. A plain acceptance ratio therefore sits near 1.000000 however the sampling is doing, and says nothing. This column reports instead the weight collected per attempt,

\[\text{HA eff.} = \frac{\sum \text{weight of the accepted trials}} {\text{trials attempted}}\]

which reduces to the ordinary acceptance ratio when every weight is 1. Two ensembles that both accept every wire-fencing trial can differ by a factor of four here, and that difference is the one that matters. The high-acceptance moves are excluded from the shooting column for the same reason; an ensemble that runs only wire fencing shows n/a there.

Exchange answers the question the swap column cannot on most ensembles. Under infinite swapping only the zero swap is an accept/reject move; every other ensemble exchanges by having its occupancy redistributed over the live paths. A cycle in which one path holds the whole ensemble exchanged nothing; a cycle split 0.9 / 0.1 exchanged a tenth. This column averages 1 - (largest share) over the analysed cycles, so 0 means the ensembles are not communicating and larger means more exchange. It is the same quantity a swap acceptance ratio measures, on the routes where no swap is ever proposed.

A ratio is shown as n/a when it cannot be formed: the move was never attempted in that ensemble (a per-interface TIS run attempts no swaps, so its swap ratio is always n/a). A move that was not attempted has no acceptance ratio – it is not a ratio of one.

Comparing these numbers with PyRETIS 3

The column is called Shoot acc. ratio in both versions but it does not count the same moves, so the two are not directly comparable.

PyRETIS 3 divided everything that was not a swap by its attempts. That denominator included the time-reversal move, and it included the null move, which is accepted by definition. PyRETIS 4 counts only the moves that actually shoot – sh, bias and wt – so it answers “how often does shooting succeed here?” rather than “how often did this ensemble keep a path?”.

The difference is not cosmetic, and it is largest exactly where it misleads most. In a 2000-cycle RETIS run of the 1D double well, the two versions report this for the outermost ensemble [3^+]:

Counted over

[3^+]

Agreement

PyRETIS 3, as reported

0.522

–

PyRETIS 3, null moves removed

0.278

–

PyRETIS 4, as reported (sh only)

0.215

–

PyRETIS 4, counted as sh + tr

0.283

0.006 from PyRETIS 3

PyRETIS 3’s 0.522 is almost twice its own shooting performance: 511 of its attempts in that ensemble were null moves, every one of them accepted. Once those are removed the two versions agree to 0.006. The same holds for the other ensembles, where null moves are rarer and the two numbers were never far apart to begin with.

So if you are reproducing a PyRETIS 3 result, expect this column to drop, most visibly in the outer ensembles – the sampling has not got worse, the ratio has stopped counting moves that could not fail.

The swap column differs for a structural reason rather than a book-keeping one: PyRETIS 3 proposed a discrete swap between every adjacent pair and could report a ratio for each, while under infinite swapping only the zero swap is an accept/reject move. Every other ensemble therefore shows n/a there and reports Exchange instead, as described above.

Older runs are still analysed. Before moves.txt the attempted moves were written into pathensemble.txt itself with a trailing Trial column, and before that into a separate trials.txt; both are read. Only one case genuinely has no answer: output written by the scheduler between those conventions, which holds accepted paths only. That is reported as n/a rather than as the 1.000000 those rows would otherwise produce.

Note

In pathensemble.txt the Mc column is a different thing: it names how the path was originally generated, which for a path that arrived by a swap is the move that created it in its previous ensemble. Read moves.txt for what was attempted here.

The same numbers are logged at the end of a run, broken down per move:

Acceptance (attempted trials, per ensemble):
  000: 34/50 = 0.6800   [sh: 34/50]
  001: 25/50 = 0.5000   [sh: 21/40 tr: 4/10]

A very low ratio usually means the interfaces are placed too far apart, or the shooting move is too aggressive for the system.