The PyRETIS analysis application¶
The PyRETIS analysis application, pyretis analyse, is used to
analyse the results from PyRETIS simulations.
See also
Simulation outputs describes the files a run writes and how to read the report this application produces – including how to tell a converged simulation from one that has not sampled enough. Calculating the rate explains how the rate constant itself is put together, and how the routes differ between TIS, RETIS, infinite swapping and the (RE)PPTIS family.
The general syntax for executing is:
pyretis analyse [-h] [-i INPUT] [-V]
where INPUT is the input file for the analysis. This file is
similar to the input file to the PyRETIS
application with some differences:
It contains the actual number of cycles/steps completed, specified by the endcycle keyword.
It contains the directory from where the simulation was executed, specified by the exe_path keyword.
It contains the number of particles in the simulation, specified by the npart keyword.
It contains settings for the analysis, defined via the analysis section.
Note
The analysis reads the SAME input file the simulation ran
from (pyretis analyse -i retis.toml), or the run’s own
output.toml (pyretis analyse -i output.toml), which
records that input with every default resolved: the analysis
takes the input sections from it and the run directory it is run
in (see Calculating the rate).
Input arguments¶
Argument |
Description |
|---|---|
-h, –help |
Show the help message and exit. |
-i INPUT, –input INPUT |
Location of the input file containing analysis settings. |
-V, –version |
Show the version number and exit. |
-skipb N, –skip-begin N |
Discard the first |
-skipe N, –skip-end N |
Discard the last |
Choosing which cycles to analyse¶
By default every sampled cycle is analysed. Two optional arguments narrow that window:
pyretis analyse -i input.toml -skipb 500
pyretis analyse -i input.toml -skipb 500 -skipe 100
The first discards the first 500 cycles – the usual reason being that the simulation had not yet equilibrated. The second additionally discards the last 100, which is what you want when a run was cut short (a wall-time kill mid-cycle) or when you are comparing two runs over the same window.
Both are counted in cycles, they can be used together or on their own, and they override skip_initial_cycles and skip_final_cycles in the input file, so you can try several windows without editing anything. The analysis log states which cycles it dropped.
The chosen window applies to everything the analysis reports: the crossing probabilities, the rate, and the acceptance ratios below.
Independently of the window, each ensemble’s statistics start at its
first accepted, non-initialisation path. The paths an ensemble holds
before that are the ones initiation put there (ld, ki or is)
and they are not samples of the ensemble, so counting them would bias
the crossing probability towards whatever the initial path happened to
do. An ensemble that has not yet been sampled at all is reported as
having no accepted paths to analyse rather than given a number built
from its initial path – if you see that, the run was too short to
reach that ensemble.
The same holds for a path marked hc: an ensemble held it at a restart
that changed [tis] high_accept, so it was accepted under the other
acceptance rule. Its rows from the restart on are left out, and the
ensemble counts again from its first path accepted after the restart; its
rows from before the restart stay in.
The WHAM analysis, and every tool that builds the combined path-data
matrix from the per-ensemble output (the training set and the interface
optimiser read the same matrix), leaves out the rows of the
initialisation paths and of the paths marked hc by the same rule.
Its skip counts paths, the initialisation paths included:
-skipb 500 discards the first 500 paths of the run by path number,
among the paths the ensembles do not hold at the end of the run, and the
rows of the initialisation paths among the paths that remain are left
out. A run whose initialisation paths all fall within the skipped paths
gives the matrix of every row after that skip.
A literal infswap_data.txt or infretis_data.txt in the run
directory is read in place of the per-ensemble output. Its rows record
no move, so every row of it is kept, the rows of the initialisation
paths included; its skip discards its first rows. The merge of several
runs (pyretis.analysis.combine_data)
writes <out>.txt, combo.txt by default, in this layout, and its
interfaces to <out>.toml. Copied into a directory as
infswap_data.txt, with <out>.toml as output.toml, the merge
is read in this way. The merge reads the paths of each run as the
analysis does, and its skip counts them in the same way.
Stone skipping without high acceptance ([tis] high_accept = false)
samples each path with weight 1. When the rows of such an ensemble carry
the crossing count of each path as HA-weight, and the counts differ from
row to row, every weighted average over them is biased, so the analysis
leaves those cycles out and its log names them. The run’s
output.toml records, as [current] ss_weight_from_cycle, the
first cycle whose rows carry the weight of the rule each path was sampled
with, and, under [current] ss_weight_before, the move list and layout
the earlier rows were written with and the paths held at that cycle. The
rows before it are checked, and so is every row of a run whose
output.toml does not record it.
Where no weight decides where a path goes (tis, pptis,
repptis), only an affected ensemble loses those cycles. In a
retis run the weights also decided, through the swaps, where each
path went, so every ensemble loses them, up to the last cycle in which a
path held at the boundary is still held. The cycles after that start
from the paths the earlier weights left behind, as a run starts from its
initial paths; give them an equilibration skip (-skipb) of their own
if they need one. An ensemble with no cycle left is refused with an
error; continuing the run with
restart adds cycles that can
be analysed. The WHAM analysis, and every tool that reads the combined
path-data matrix, sums every cycle of a path, so it refuses such a run.
Two cases leave no trace in the run’s files: a change of
[tis] high_accept that output.toml does not record (a state
file edited by hand and resumed from), and rows appended after the
boundary by a version of PyRETIS that writes the crossing count.
Acceptance ratios¶
The report’s Pathensemble data table gives, per ensemble, the fraction of attempted moves that were accepted – separately for the shooting-like moves and for the swapping moves.
These are counted from moves.txt – one line per attempted
move: the cycle, the move that ran, and the status it ended in (ACC,
or a rejection code such as BWI or NCR; see
Path types and rejection reasons). It is the only record of a
rejection. pathensemble.txt beside it says which path each
ensemble holds, so every one of its lines is an accepted path by
construction – counting those would report every move as accepted. See
moves.txt – what the sampler attempted for the file’s columns.
A cycle in which two ensembles were picked together is a swap, whatever
shooting move those ensembles are configured with, and is recorded in
both partners: under the + code in the lower-index one and the -
code in the higher. The pair of codes says which swap it was –
s+/s- for the zero swap between [0^-] and [0^+],
p+/p- for the PPTIS/REPPTIS swap between adjacent ensembles.
That is what makes the swap ratio computable at all.
Two further columns report what counting acceptances cannot.
HA eff. is for the high-acceptance moves – wire fencing and stone
skipping. Those are built so that a trial is accepted almost every time;
what varies is the weight the accepted path carries, its crossing
count. A plain acceptance ratio therefore sits near 1.000000 however
the sampling is doing, and says nothing. This column reports instead the
weight collected per attempt,
which reduces to the ordinary acceptance ratio when every weight is 1.
Two ensembles that both accept every wire-fencing trial can differ by a
factor of four here, and that difference is the one that matters. The
high-acceptance moves are excluded from the shooting column for the same
reason; an ensemble that runs only wire fencing shows n/a there.
Exchange answers the question the swap column cannot on most
ensembles. Under infinite swapping only the zero swap is an
accept/reject move; every other ensemble exchanges by having its
occupancy redistributed over the live paths. A cycle in which one path
holds the whole ensemble exchanged nothing; a cycle split 0.9 / 0.1
exchanged a tenth. This column averages 1 - (largest share) over the
analysed cycles, so 0 means the ensembles are not communicating and
larger means more exchange. It is the same quantity a swap acceptance
ratio measures, on the routes where no swap is ever proposed.
A ratio is shown as n/a when it cannot be formed: the move was never
attempted in that ensemble (a per-interface TIS run attempts no swaps, so
its swap ratio is always n/a). A move that was not attempted has no
acceptance ratio – it is not a ratio of one.
Comparing these numbers with PyRETIS 3¶
The column is called Shoot acc. ratio in both versions but it does not count the same moves, so the two are not directly comparable.
PyRETIS 3 divided everything that was not a swap by its attempts. That
denominator included the time-reversal move, and it included the null
move, which is accepted by definition. PyRETIS 4 counts only the moves
that actually shoot – sh, bias and wt – so it answers “how
often does shooting succeed here?” rather than “how often did this
ensemble keep a path?”.
The difference is not cosmetic, and it is largest exactly where it
misleads most. In a 2000-cycle RETIS run of the 1D double well, the two
versions report this for the outermost ensemble [3^+]:
Counted over |
|
Agreement |
|---|---|---|
PyRETIS 3, as reported |
0.522 |
– |
PyRETIS 3, null moves removed |
0.278 |
– |
PyRETIS 4, as reported ( |
0.215 |
– |
PyRETIS 4, counted as |
0.283 |
0.006 from PyRETIS 3 |
PyRETIS 3’s 0.522 is almost twice its own shooting performance: 511 of its attempts in that ensemble were null moves, every one of them accepted. Once those are removed the two versions agree to 0.006. The same holds for the other ensembles, where null moves are rarer and the two numbers were never far apart to begin with.
So if you are reproducing a PyRETIS 3 result, expect this column to drop, most visibly in the outer ensembles – the sampling has not got worse, the ratio has stopped counting moves that could not fail.
The swap column differs for a structural reason rather than a
book-keeping one: PyRETIS 3 proposed a discrete swap between every
adjacent pair and could report a ratio for each, while under infinite
swapping only the zero swap is an accept/reject move. Every other
ensemble therefore shows n/a there and reports Exchange instead,
as described above.
Older runs are still analysed. Before moves.txt the attempted
moves were written into pathensemble.txt itself with a trailing
Trial column, and before that into a separate trials.txt;
both are read.
Only one case genuinely has no answer: output written by the scheduler
between those conventions, which holds accepted paths only. That is
reported as n/a rather than as the 1.000000 those rows would
otherwise produce.
Note
In pathensemble.txt the Mc column is a different thing:
it names how the path was originally generated, which for a path
that arrived by a swap is the move that created it in its previous
ensemble. Read moves.txt for what was attempted here.
The same numbers are logged at the end of a run, broken down per move:
Acceptance (attempted trials, per ensemble):
000: 34/50 = 0.6800 [sh: 34/50]
001: 25/50 = 0.5000 [sh: 21/40 tr: 4/10]
A very low ratio usually means the interfaces are placed too far apart, or the shooting move is too aggressive for the system.