Benchmark issues

Hello everyone,

Hope you’re doing great. I have been trying to justify that I understand ASPECT a little bit by doing benchmark against some papers with my own plugins. However, it turns out that I probably don’t understand it at all. Here’s the list of papers.

[1] Laneuville et al., 2013 (Asymmetric thermal evolution of the Moon): failed to reproduce because some details are not clear, esp. the lower limit of viscosity. The authors used GAIA. Laneuville et al., 2018 (Distribution of radioactive heat sources and thermal history of the Moon) used same model, but the inner core growth in main text is not consistent with figures. So it’s impossible to reproduce it.

[2] Zhang et al., 2013 (A 3‐D numerical study of the thermal evolution of the Moon after cumulate mantle overturn: The importance of rheology and core solidification): failed to reproduce because numbers in main text are not consistent with figures, which is observed after a couple of tests and also confirmed from the authors. The authors used CitComS.

[3] Euen et al., 2023 (A comparison of 3-D spherical shell thermal convection results at low to moderate Rayleigh number using ASPECT (version 2.2.0) and CitComS (version 3.3.1)): precisely reproduced with my own plugins. But this paper shows simple cases.

[4] Citron et al., (Effects of Heat‐Producing Elements on the Stability of Deep Mantle Thermochemical Piles): successfully reproduced without my own plugins. The authors used ASPECT also.

[5] Siqi Zhang et al., 2016 (The early geodynamic evolution of Mars-type planets): still ongoing, but the detail of initial perturbations of temperature and lithosphere thickness are not given in the main text, which is not so good.

The failure of reproduction of these papers are partially due to (1) unspecified important details, (2) inconsistencies in papers itself, (3) systematic differences among different code for complex models, and (4) misunderstanding of papers and/or ASPECT. (1-2) cannot be avoided easily; To eliminate systematic differences, I moved from papers [1-2] (GAIA and CitComS) to [3-5] (ASPECT). I checked all the ASPECT publications to select papers that (a) have lines instead of contours, lines can be faithfully reproduced; (b) global thermo-chemical evolution instead of plate tectonics which is the major part of the publication list; (c) intermediate difficulty. Papers [3-5] are thus found.

I have been doing the benchmark for almost a year and still cannot reproduce any papers like Laneuville 2013 or Zhang 2013, 2016. This is quite frustrating. So I’m wondering if any of ASPECT developers or experienced users can give me some advice. I’m not expecting at all that anyone would spend time on reading my plugins, but it would be fantastic if anyone could do that. Additionally can we add a tag to the publication in the list to show if the paper is successfully benchmarked by other users, other ASPECT versions or other codes, so that new comers can follow?

Regards,
Mingming

Hi Mingming,

I don’t have much experience reproducing papers using ASPECT, but I do have a lot of experience benchmarking and testing other codes and models, so perhaps I can help in that respect.

My experience with the academic literature (including my own papers) suggests there are mistakes and typos in most modelling papers; sometimes in the code, but more often in the descriptions. For this reason, I find the following invaluable:

  1. access to the original source code and input files
  2. ability to correspond with the original authors
  3. an understanding of how close I am likely to get to the published result with a different code (i.e., is this a good benchmark)
  4. an understanding of how far away I might end up from the published result in the absence of any bugs (i.e. are there aspects of the study which make it a poor choice for a benchmark).

It seems to me that your struggles so far are related to a lack of one or more of these. I fully appreciate the frustration that you are feeling, and I fear it might be quite difficult to succeed in your goals without these things.

All of this is somewhat removed from the auxiliary goal of understanding ASPECT better, a goal with which we might be able to help more effectively!

Best wishes,
Bob

Hi Bob,

Thanks for your comments. Yes, I understand that there could be more or less typos or mistakes in papers, and sometimes I can correct them. But sometimes a series of typos or mistakes can never be easily corrected. When this happens, I usually make contact with the authors, as you suggested. Unfortunately, some authors themselves cannot understand or remember what they did to these numbers, like the ones I mentioned in [1-2]. I’m not judging people and I understand typos and mistakes cannot be escaped completely, but it’s the numbers that convince people, that tell a story, that is what numerical modelling is all about.

Yes, reading source code and input file is the best way to understand what’s been done in the paper. And this is what I’ve been doing recently.

Anyway, thanks for sharing your experience.

Regards,
Mingming

I’ll second Bob here: I have spent enormous amounts of time trying to replicate studies of all kind, and I have rarely been truly successful (with my own codes or with ASPECT). Many times I had to contact the author(s) to clarify things and/or found issues in the article itself (wrong parameters reported, wrong units, caption that do not correspond to the figure, …).

Best,

Cedric

This is frustrating, and this would be a big gap that new comers like me cannot handle and may fail to continue. So how people justify what’s been done in their publications is correct?

Unfortunately, I’m working a supervisor who does not know much about ASPECT/fluid dynamics, and he expects a faithful/precise reproduction of some complex models. I can precisely reproduce some benchmark papers with my plugins, but cannot reproduce complex models. I cannot persuade him and myself that my plugins are correctly built. We’re running into a new paper benchmark – failure – new paper benchmark – failure mode, we’re losing faith on each other and the project. God please save us!

Hi Mingming,

So how people justify what’s been done in their publications is correct?

The ASPECT developers try hard to publish everything alongside their papers (normally with a Zenodo DOI). This includes the exact version of the software, the input files, and key output. For all of our papers where we follow this strategy, you should be able to perfectly reproduce our model results in ASPECT.

This is not standard in our discipline - although model and data publishing practises have generally improved over the last 10 years, many academic groups do not publish their code or input files. Reasons for this include academic protectionism, a lack of knowledge of public storage methods and versioning standards, insecurity and laziness (this list is not meant to be exhaustive!). Basically, being human.

I’m working with a supervisor who does not know much about ASPECT/fluid dynamics, and he expects a faithful/precise reproduction of some complex models.

Is this your eventual goal, or is there a future scientific question you want to address, or a paper you want to write?

I can precisely reproduce some benchmark papers with my plugins, but cannot reproduce complex models.

This is normal in geodynamics. Most complex problems are very sensitive to the exact formulation, initial conditions and numerical noise. I would not expect you to be able to reproduce complex models perfectly. A better, more tractable goal is to be able to reproduce existing benchmarks which are meant to have a well-defined solution that is insensitive to precise setup. But there are not many benchmarks in the literature - even some problems that are called “benchmarks” are particularly poorly suited to being code-agnostic banchmarks.

we’re losing faith on each other and the project

I hope that this can be rectified. We’re happy to provide advice and guidance (within reason!). My view is that real benchmarking from scratch with geodynamics software can be a very hard thing to do (Cedric and I have decades of experience between us, and we still get stuck). I think if you want to gain experience with ASPECT, it might be more efficient to look at a few of the benchmarks and cookbooks that we have provided.

God please save us!

I don’t believe God responds to this chat, you’ll have to make do with humans :wink:

Best wishes,
Bob

It’s perhaps worth adding that the inability to reproduce what others have done is a widespread (and frustrating) issue in the sciences. You may want to take a look at Replication crisis - Wikipedia . The length of that article is perhaps an indicator for how much of an issue this is in other disciplines. For a specific project that tried to replicate highly-cited papers, see also here: Reproducibility Project - Wikipedia .

There is no question that all of that is (i) frustrating, and (ii) not actually a good advertisement for the scientific method. In your specific case, in addition to slowing your progress, it also puts into question our own field, something many of us will feel strongly about on a personal level. I think a perspective you might want to think about is whether for the papers you cannot exactly reproduce, the fundamental conclusions of the paper are wrong, or whether the paper’s message still holds even though perhaps your simulations do not exactly match those of the paper. My personal perspective is that in all disciplines, a substantial fraction of papers is not exactly reproducible, but that at least in relatively slow-moving fields like ours, the number of papers with conclusions that are fundamentally wrong or invented is relatively small. I don’t have concrete evidence for the second half of this claim, but would be interested in hearing what your take is for the papers you can’t reproduce the concrete simulations.

Best
WB

Best
WB

Hi Bob,

First of all, thank you so much for wasting your time on this naive question, I deeply appreciate that.

Yes, benchmarks and cookbooks in ASPECT manual is the first place to look at. In fact, I have read the manual a few times and also run models correspondingly. That’s why I can write, compile and run my own plugins, though not all my plugins work.

I also understand that most numerical modelers are trying their best to be open and rigorous, but sometimes sh*t happens. It’s just a matter of being lucky or not.

Anyway, I’ll keep on it.

Cheers
Mingming

Hi Bangerth,

Nice to hear from you! I guess what you’ve just showed us is basically true and kinda frustrating. The value that people can appreciate from a paper depends on what people are looking for, but a paper generally gives us at least 3 aspects to evaluate: idea, method, and conclusion. Some papers present brilliant ideas, even the method is simple; some papers show elegant and complete method that people are after, so that the modeling is more accurate or realistic; some may have strong, meaningful and general conclusions.

For senior researchers, they care more about the ideas and conclusions because they are experienced numerical modelers.

For freshmen, they focus more on how the problem is solved, i.e., the method. This is where I’m. I believe (1) provided all the info in a paper is correct and complete, people should reproduce the paper up to floating point precision; (2) people are trying their best to be informative and rigorous, and they don’t make up intentionally. For papers that I cannot precisely reproduce, in most cases, they don’t show consistent and complete info. Obviously only when given complete setups can we determine whether our customized plugins are correctly formulated. But still, I’d consider the conclusion is meaningful and reasonable for non reproducible papers in general. Sometimes even if the conclusion is unphysical or unreasonable, I can still appreciate its value, because people make mistakes and science is pushed forward this way.

In terms of reproducing papers, I just focus on what I’m told in the paper, even when the physics breaks, I’ll just follow it. I’m trying to be a machine when I reproduce a paper.

Thanks for spending time on this issue.

Cheers,
Mingming

Hi @cedrict @bangerth @bobmyhill ,

I have successfully installed ASPECT-1.4-pre which is the branch used for the paper I have been trying to benchmark against. An input file is also provided in the source code. However, the input file is not the one used for the paper, it’s merely a test of an even older paper, a lot of parameters are missing. I finally realized why I always failed, it’s this info difference that makes me so struggling. Only when I’m a member of the group that published the paper, can I possibly reproduce the results. I have no idea what to do next except waiting for response from the authors. Any suggestions?

Best,
Mingming

@optimux There isn’t very much you can do. If they respond then great, if they don’t, then you can’t force them. I think the way you should see that is that your inability to reproduce these results isn’t your fault. It’s theirs for not describing their methods well enough, for not archiving their scripts, and for not getting back to you about your questions. You’re not going to change that, though, and so the only reasonable decision for you is to move on – do your own stuff. In doing so, keep in mind that someone in the future might ask you the same questions, so make sure you’re doing a better job at documenting and archiving in your own publications!

Best
WB

@optimux Thank you for raising this topic. During the first two years of my graduate studies, I also spent much of my time benchmarking and reproducing published results, and I am still doing similar work recently. In my experience, unsuccessful reproduction is quite common. It is more often caused by an incomplete understanding of the paper or numerical method, undocumented implementation details, different parameter and post-processing conventions, or differences between codes and versions, rather than serious errors by the authors.

Recently, I found a persistent discrepancy while reproducing a semi-analytical solution from a classic paper published about 20 years ago. After I contacted the authors, they confirmed that the published equations contained a few minor typos. However, I have also sent many inquiries and received replies to only a small number of them, while most went unanswered, so I fully understand the frustration.

I believe that reproduction is a scientific investigation in its own right. Even when a result cannot be reproduced completely, carefully investigating the problem can still teach us a great deal.

@optimux - Adding onto to what others have said in this discussion, reproducing the results of prior publications can be very difficult.

In fact, obtaining equivalent results between codes is challenging even when the different groups are working directly together on a well defined problem.

For example, I think you would appreciate the findings, discussion, and recommendations from Buiter et al. (2016), which focused on benchmarking brittle wedge thrust experiments between different numerical codes and analogue experiments. Due to the different results observed between codes, the authors list a minimum of 16 parameters (physical and numerical) that should be included with every geodynamic study using plasticity.

Having worked with many codes over the years and accordingly attempted to reproduce prior studies, my general finding has been that the results of many studies are reproducible in the sense that the key observations and processes can be replicated, but often over a slightly different parameter space. These differences can reflect the underlying numerics (solver schemes, tolerances, discretization, time stepping, etc, etc, etc), slight differences in how the underlying physics are implemented (averaging rheology over different fields, averaging, etc, etc) or as others have noted small typos.

As @tiannh7 noted, even if you can’t exactly reproduce the problem in question you will still learn quite a bit about the codes work in the process and then apply this to your specific investigation. I hope this is helpful in determining when you will be comfortable moving on from the benchmarking stage in your project.

Hi @jbnaliboff @tiannh7 @bangerth ,

Thanks for sharing your experience and suggestions.

The reasons why I still insist on a precise reproduction include my reproduction experience and a requirement from my present supervisor. I have successfully reproduced a 1990s paper on alloy macrosegregation during cooling (similar to chemical differentiation in silicate melt) from scratch, absolute scratch (the authors have passed away a long time ago, no code can be referenced and no question can be sent; besides I was in an institute of geochemistry, nobody knows about fluid dynamics). It’s a fully coupled thermo-chemical N-S system. This reproduction strongly built up a belief that numerical modeling papers are reliable and can be reproduced. Later I successfully reproduced thermal convection in an annulus using spectral method which I have never used before, so this is also a fresh reproduction. Recently, I reproduced a two-phase flow paper using FEniCSx. But now when it comes to thermo-chemical evolution of terrestrial planets, I found that their models are so fragile and easy to lose necessary info.

My supervisor has an excellent idea that requires thermo-chemical evolution of the Moon. So he asked to benchmark some papers and then apply to the Moon. He is not in the community of geodynamics but planetary impacts. It’s the fact that he does not know much about the ecosystem of geodynamics that forces him to ask for a precise reproduction of prior papers. When some differences are present, he believes that something is wrong on my side, even when sometimes the paper itself shows inconsistency. I cannot persuade him without showing matched lines.

I guess the only way to make through is either to ask someone to verify my code or have in-depth discuss with the authors. I’ll keep waiting for the authors for a week or so, if this fails, I guess I need to fly there. This is my last chance, otherwise I have to quit.

Regards,
MM

@optimux If you believe that your reproduction attempts are set up correctly but that you need feedback from paper authors to figure out details, and if these authors do not or cannot help you, then I think that your reproduction is simply not feasible. I understand that that puts you in a tough spot with your adviser. Your only way out is to have that conversation with your adviser – it isn’t your fault, there just isn’t anything you can do about it. If it helps you with that discussion, point your adviser to the conversation here.

Best
WB

Hi @bangerth ,

Thanks for your advice, that sounds practical. By the way, is it possible to change this? This is not fair for freshmen and not good for science community.

Best,
MM

@optimux All here seem to agree that this is an awkward situation for you, and not great for science in general. But in the end, “science” is a collection of people. We need to change people, but that’s hard. The only thing we can do is change ourselves and the ones around us, by reflecting on what we do and by having conversations with friends and colleagues. That’s all we can really do: Change is local. But that’s also what we’re doing in this thread.

Best
WB

1 Like