Table of Contents
Fetching ...

Very-Long Baseline Interferometry Imaging with Closure Invariants using Conditional Image Diffusion

Samuel Lai, Nithyanandan Thyagarajan, O. Ivy Wong, Foivos Diakogiannis

TL;DR

This work tackles the ill-posed inverse problem of reconstructing VLBI source images from calibration-free closure invariants by introducing GenDIReCT, a conditional diffusion-based generative pipeline coupled with a CNN compressor. By training a latent diffusion model on a CIFAR-10–augmented dataset and conditioning on closure invariants, the approach yields a distribution of plausible reconstructions that are refined to a single image via a data-aware compression step, achieving high fidelity (ρ_{ m NX} ≳ 0.9) and good data adherence (χ^2_{ m CI} ≲ 1) across trained and several untrained morphologies and noise levels. The method demonstrates competitive performance on ngEHT challenge data and offers a calibration-independent, reproducible imaging framework with minimal hyperparameter tuning. The results highlight the potential of closure invariants for robust VLBI imaging and point to future extensions in dynamic, polarimetric, and multi-frequency VLBI analyses, including real data applications such as M87.

Abstract

Image reconstruction in very-long baseline interferometry operates under severely sparse aperture coverage with calibration challenges from both the participating instruments and propagation medium, which introduce the risk of biases and artefacts. Interferometric closure invariants offers calibration-independent information on the true source morphology, but the inverse transformation from closure invariants to the source intensity distribution is an ill-posed problem. In this work, we present a generative deep learning approach to tackle the inverse problem of directly reconstructing images from their observed closure invariants. Trained in a supervised manner with simple shapes and the CIFAR-10 dataset, the resulting trained model achieves reduced chi-square data adherence scores of $χ^2_{\rm CI} \lesssim 1$ and maximum normalised cross-correlation image fidelity scores of $ρ_{\rm NX} > 0.9$ on tests of both trained and untrained morphologies, where $ρ_{\rm NX}=1$ denotes a perfect reconstruction. We also adapt our model for the Next Generation Event Horizon Telescope total intensity analysis challenge. Our results on quantitative metrics are competitive to other state-of-the-art image reconstruction algorithms. As an algorithm that does not require finely hand-tuned hyperparameters, this method offers a relatively simple and reproducible calibration-independent imaging solution for very-long baseline interferometry, which ultimately enhances the reliability of sparse VLBI imaging results.

Very-Long Baseline Interferometry Imaging with Closure Invariants using Conditional Image Diffusion

TL;DR

This work tackles the ill-posed inverse problem of reconstructing VLBI source images from calibration-free closure invariants by introducing GenDIReCT, a conditional diffusion-based generative pipeline coupled with a CNN compressor. By training a latent diffusion model on a CIFAR-10–augmented dataset and conditioning on closure invariants, the approach yields a distribution of plausible reconstructions that are refined to a single image via a data-aware compression step, achieving high fidelity (ρ_{ m NX} ≳ 0.9) and good data adherence (χ^2_{ m CI} ≲ 1) across trained and several untrained morphologies and noise levels. The method demonstrates competitive performance on ngEHT challenge data and offers a calibration-independent, reproducible imaging framework with minimal hyperparameter tuning. The results highlight the potential of closure invariants for robust VLBI imaging and point to future extensions in dynamic, polarimetric, and multi-frequency VLBI analyses, including real data applications such as M87.

Abstract

Image reconstruction in very-long baseline interferometry operates under severely sparse aperture coverage with calibration challenges from both the participating instruments and propagation medium, which introduce the risk of biases and artefacts. Interferometric closure invariants offers calibration-independent information on the true source morphology, but the inverse transformation from closure invariants to the source intensity distribution is an ill-posed problem. In this work, we present a generative deep learning approach to tackle the inverse problem of directly reconstructing images from their observed closure invariants. Trained in a supervised manner with simple shapes and the CIFAR-10 dataset, the resulting trained model achieves reduced chi-square data adherence scores of and maximum normalised cross-correlation image fidelity scores of on tests of both trained and untrained morphologies, where denotes a perfect reconstruction. We also adapt our model for the Next Generation Event Horizon Telescope total intensity analysis challenge. Our results on quantitative metrics are competitive to other state-of-the-art image reconstruction algorithms. As an algorithm that does not require finely hand-tuned hyperparameters, this method offers a relatively simple and reproducible calibration-independent imaging solution for very-long baseline interferometry, which ultimately enhances the reliability of sparse VLBI imaging results.
Paper Structure (25 sections, 11 equations, 8 figures, 2 tables)

This paper contains 25 sections, 11 equations, 8 figures, 2 tables.

Figures (8)

  • Figure 1: Diagram of the GenDIReCT architecture. In the diffusion network, closure invariants measured for images in the training dataset are used to condition the denoising UNet and furthermore, as targets for the convolutional network. The conditional UNet is trained to reverse the diffusion process by taking gradient descent steps on the $\mathcal{L}_{\textsc{GenDIReCT}}$ objective function, defined in Equation (\ref{['eq:gendirect-loss']}) The result is a sample of images, $p_\theta(x|y)\sim p_{\rm data}(x|y)$, which are inputs for the convolutional neural network. The CNN learns the optimal compression for the set of the sampled images on the image axis based on the loss function defined in Equation (\ref{['eq:cnn-loss']}), which ensures that the final reconstructed image is consistent with the input closure invariants.
  • Figure 2: Illustration of the output from each layer of the GenDIReCT architecture and typical imaging procedure. The illustration utilises a simulated image of Sgr A*, which is not part of the training dataset. Random noise is used to initialise the denoising UNet conditioned on closure invariants observed from the ground truth image to create a sample of latent information, which can be decoded into images. We show 64 images sampled from the diffusion model, all of which show a crescent-like structure. The final image is reconstructed by the CNN by learning the optimal compression of the diffusion sample.
  • Figure 3: Plot of the maximum normalised cross-correlation image fidelity metric, $\rho_{\rm NX}$ (dots), of the final reconstructed image and the relative CRPS (crosses) of the diffusion output as a function of the closure invariants' signal-to-noise ratio on a simulated observation of Sgr A*. Vertical dotted lines mark SNR thresholds of 31.4, 10, and 3, which correspond to the median phase calibrated Stokes I component SNR of the primary M87 EHT dataset EHT_2019_Data, a SNR threshold for self-calibration, and a threshold commonly used for low-SNR flagging, respectively.
  • Figure 4: Results of the GenDIReCT image reconstruction pipeline on a variety of test images, where the first five basic shapes (Gauss, Double, Ellipse, Ring, and Crescent) are represented in the training dataset, but the latter three ($m$-ring, Centaurus A, and Einstein) are examples of untrained morphologies. The first row presents the ground truth image from which visibilities and subsequently closure invariants are derived. The GenDIReCT middle row presents the final reconstruction from this work's imaging pipeline, alongside the maximum normalised cross-correlation $\rho_{\rm NX}$ image fidelity metric. The bottom row displays the ground truth closure invariants as black diamonds and reconstruction closure invariants as grey points. The final $\chi^2_{\rm CI}$ goodness-of-fit metric is shown. The Einstein model is illustrated with an inverted colormap for enhanced visual clarity.
  • Figure 5: From left to right: Ground truth image, median image reconstruction, median absolute deviation image of all reconstructions, and ratio image of the median to the median absolute deviation, which illustrates a 'signal-to-noise' ratio of image reconstructions. Large values in the ratio image indicate pixels with low variance relative to the mean pixel intensity. Contours on the ratio image highlight pixels with high perceptual hash variance, which corresponds to higher morphological uncertainty. They are observed to occur on regions of low signal-to-noise' ratio, and thus do not signify an appreciable morphological difference.
  • ...and 3 more figures