diff --git a/UseCases.tex b/UseCases.tex index 0fcd182..e9ae626 100644 --- a/UseCases.tex +++ b/UseCases.tex @@ -1,15 +1,72 @@ -\subsection{Event-List Data and Responses} +Here are some science cases collected from various astronomical projects for this note to represent typical data discovery scenarios in high energy physics. +Each use case presents one strategy for building a ADQL query for an extended TAP service, based on the original {\tt ivoa.obscore} table and on the {\tt ivoa.obscore-hea} table extension. + +We distinguished three categories of datasets in this note: +primary observation datasets (typically event-lists and event-bundles for \gls{HEA}), response function datasets, and advanced data product datasets. +In agreement with Semantics WG we have 3 existing/draft/proposed vocabularies: +\begin{itemize} +\item \url{http://www.ivoa.net/rdf/product-type} used for data products served in {\tt ivoa.obscore} table; +\item \url{http://www.ivoa.net/rdf/response-type} to expose response data products; +\item an Advanced data product type vocabulary currently discussed here: \\ +{\small \url{https://github.com/ivoa-std/VEPs/pull/21/changes/c4195e1975341881d8fda22cb851ff11b5b74c1f#top}}.\\ +\vspace{-20pt} +\item[] Note that the Radio astronomical community also requires the definition of new types of data products for their science ready datasets as ``advanced dataset''. +\end{itemize} + +Three strategies have been considered in this note to organize data discovery based on an extension of ObsCore 1.1 in TAP\null. Each of them needs to be evaluated in various situations and explored further in terms of clarity for the user and robustness. + +\begin{itemize} +\item {\bf Proposal 1}: Serve all data products in the {\tt ivoa.obscore} table. The attribute {\em dataproduct\_type\/} in {\tt ivoa.obscore} can contain any term from one of the three vocabularies \textsl{product-type}, \textsl{advanced-product-type}, or \textsl{response-type}: +All products are distributed via the {\tt ivoa.obscore} table as in +Chandra use cases A.1.8, A.2.1, and A.2.3, CTAO use case A1.9 and A2.4, and SWGO use case A1.3. + +This has the advantage that it is easy for the end-user to understand. All data products are accessed from the same table and the user does not have to know details about which data products need to be queried from what table. + +A potential disadvantage is that the results of queries for multiple types of datasets will be harder to correlate, especially if the different data products types don't have a one-to-one cardinality relationship. For example, if you search for {\bf hea-event-list} or {\bf light-curve} in a sky region together with {\bf response-function} files, for instance {\bf psf}s, the query response may provide all datasets together mixing the various data product types together. The user will therefore have to determine which {\bf psf}(s) are associated with each {\bf hea-event-list} or {\bf light-curve}. + +The user could for example choose to use GROUP BY {\em obs\_id\/}, but still there may be several {\bf light-curve}s and/or {\bf psf}s within the selected region, so the issue is to organize and group properly corresponding {\bf light-curve}s with their associated {\bf psf} files, especially if the point spread function varies over the field of view. + +In practice, many advanced data products of the same type may be associated with a single {\bf hea-event-list} and so archives that serve such products will generally provide adequate metadata to allow the correlations to be performed. These metadata may not be directly accessible through the ObsCore TAP service however, which would place the onus on the end-user to perform these correlations after-the fact. + +\item {\bf Proposal 2}: Add new terms to the product-type vocabulary and distinguish among response types and advanced product types using {\em dataproduct\_subtype\/} in {\tt ivoa.obscore}. +The terms pdf, draws, region, response-function are added to product-type vocabulary. +All products are distributed with the {\tt ivoa.obscore} table. +See use cases A.1.10 and A.1.13 for KM3NeT and Chandra use case A.1.8. + +This approach is included largely for compatibility with the ObsCore Recommendation Version 1.1, which has a limited set of {\em dataproduct\_type\/}s and so requires the use of {\em dataproduct\_subtype\/} to discriminate between different data product types that share the same {\em dataproduct\_type\/}. + +A disadvantage of this approach is that queries on two ObsCore attributes are required to identify certain types of data products, for example different types of {\em response-function\/}s. The interpretation of the {\em dataproduct\_subtype\/} attribute may be ambiguous as the content could be ({\em e.g.\/}) a type of {\bf response-function}, or a category of image ({\em e.g.\/}, an excess map or significance map) or some specific description chosen by a specific archive. + +A standardized vocabulary for {\em dataproduct\_subtype\/} will be beneficial. We have proposed standardized terms for various types of {\bf response-function}s in \S~5. However, there are several types of images that are commonly used in \gls{HEA} that could, with some additional discussion, also be standardized. + +However, the ability for each individual archive to use {\em dataproduct\_subtype\/} ``to more precisely define the nature of the dataset'' (ObsCore Recommendation Version 1.1 \S~4.1) should not be discounted. Each archive is likely to have many different types of data products that would otherwise have the same {\em dataproduct\_type\/} that will need to be discriminated to be useful. Some data products may be unique to a particular observatory but nevertheless highly significant for end-users of that observatory's data. + +\item {\bf Proposal 3}: Separate the three kinds of products defined in the three distinct vocabularies as proposed with Semantics WG. +Implement one table for each kind: +\begin{itemize} +\item {\tt ivoa.obscore} uses {\em dataproduct\_type\/} from the \textsl{product-type} vocabulary; +\item {\tt ivoa.response} uses {\em dataproduct\_type\/} from the \textsl{response-type} vocabulary; +\item {\tt ivoa.adp} uses {\em dataproduct\_type\/} from the \textsl{advanced-product-type} vocabulary. +\end{itemize} + +The benefits lie in tracing clearly the connections between response files and/or advanced data products to datasets from the currently included in the \textsl{product-type} vocabulary. This separation also addressed concerns that most {\bf response-function}s are not ``on-the-sky'' data products, despite their importance to \gls{HEA} data analysis. (Advanced data products currently proposed for inclusion in ObsCore are all ``on-the-sky'' data, although the definitions of the products are sufficiently general that this is not necessarily the case.) +The disadvantage of this approach is that the end-user will need to know which vocabularies would need to be queried for which data product types and this may require a level of knowledge of IVOA implementation internals that would go beyond what is otherwise required of end-users. This will be more important for end-users who prefer to write their own queries directly using ADQL (for example, using TOPCAT), and could be particularly troublesome if the end user needs to change the syntax of their {\em dataproduct\_type\/} query depending on which vocabulary is used. + +This strategy still needs to be investigated in detail. Some hints for implementing this approach are provided in Appendix~\ref{sec:accessoptionsappendix}. +\end{itemize} + +\subsection{Event-List Data and Responses} \subsubsection{Use Case --- Search for event lists surrounding Sgr A*, for example for an X-ray morphological study} {\em Identify all \gls{HEA} event lists encompassing Sgr~A* for initial selection for subsequent X-ray morphological studies. Since the focus is on X-ray morphological studies, only the event lists and not the event bundles are desired.\/} - +%proposal regular Obscore table \medskip \noindent Find all datasets satisfying: \begin{enumerate}[(i)] \item Target name = ``Sgr A*'' or position inside 30 arcmin from (266.4168, $-29.0078$), - \item dataproduct\_type = ``event-list''. + \item dataproduct\_type = ``hea-event-list''. \end{enumerate} \begin{verbatim} @@ -26,6 +83,7 @@ \subsubsection{Use Case --- Search for event lists that include a fully calibrat {\em Identify all event lists that include the BL Lac, have a fully calibrated spectral axis (i.e., spectral responses have already been applied), and have at least 10,000 events. These data will be used to prepare slides for a presentation. Note that since calib\_status = 2 may not specify that the spectral axis is fully calibrated in physical units (HEA event lists are often considered ``calibrated'' even if the spectral axis is in pulse height units) the calibration status of the spectral axis must be checked explicitly.\/} +%proposal regular Obscore table with join \medskip \noindent Find all datasets satisfying: \begin{enumerate}[(i)] @@ -53,6 +111,7 @@ \subsubsection{Use Case --- Search for SWGO event lists and their \glspl{IRF} f {\em Identify all event lists and their associated \glspl{IRF} of the region of Cygnus loop ($3^{\circ}$ diameter) taken with SWGO. Only data of the event type `very-good' are selected, in order to limit the amount of downloaded data.\/} +%proposal 1 : all entries in ivoa.obscore with 3 vocabularies used for dataproduct_type \medskip \noindent Find all SWGO datasets satisfying: \begin{enumerate}[(i)] @@ -62,7 +121,7 @@ \subsubsection{Use Case --- Search for SWGO event lists and their \glspl{IRF} f \item event\_type = ``very-good''. \end{enumerate} -First, run the ObCore query: +First, run the ObsCore query: \begin{verbatim} SELECT * FROM ivoa.obscore NATURAL JOIN ivoa.obscore_hea @@ -74,12 +133,15 @@ \subsubsection{Use Case --- Search for SWGO event lists and their \glspl{IRF} f AND (event_type = 'very-good') \end{verbatim} -Then, for each row of the output, we identify the nature of the data product, and retrieve them using the ``access\_url''. +Then, for each row of the output, we identify the nature of the data product, and retrieve them using the {\em access\_url\/}. \subsubsection{Use Case --- Search for event bundles via DataLink that include Cas A for a TeV spectromorphology study} {\em Identify all event bundles (event lists and their associated \glspl{IRF}) that include the Cas A SNR for subsequent TeV spectromorphology studies from a VERITAS data release. Since the instrumental responses are mandatory to remove instrumental effects, the event bundles that include the \glspl{IRF} are required.\/} +%proposal regular Obscore table with product-type update for hea-event-list +%the extra files exposed in the data link depend on the data provider strategy but the vocabularies used in the content-descriptor column must be defined in one of the 3 vocabularies. + \medskip \noindent Find all VERITAS datasets satisfying: \begin{enumerate}[(i)] @@ -100,12 +162,12 @@ \subsubsection{Use Case --- Search for event bundles via DataLink that include C AND (access_format = ’application/x-votable+xml;content=datalink’) \end{verbatim} -Then, for each row of the output, we get access to a DataLink table (in VOTable) describing associated data linked to the hea-event-list dataset using the ``access\_url'' column value of the response ObsCore table. - +Then, for each row of the output, we get access to a DataLink table (in VOTable) describing associated data linked to the {\bf hea-event-list} dataset using the {\em access\_url\/}' attribute value of the ObsCore query response table. \subsubsection{Use Case --- Search for event bundles that include Cas A for X-ray spectrophotometric evolution studies} {\em Identify all event bundles that include the Cas A SNR and have at least 1 million events for subsequent spectrophotometric studies of the SNR expansion. Since only a few observations are expected to match this request and because the focus is on X-ray spectrophotometric studies, the event bundles that include the responses or the ancillary products used to make the responses are required.\/} +%proposal regular Obscore table joined to ivoa.obscore-hea \medskip \noindent Find all datasets satisfying: @@ -125,10 +187,10 @@ \subsubsection{Use Case --- Search for event bundles that include Cas A for X-ra AND (ev_xel >= 1000000) \end{verbatim} - \subsubsection{Use Case --- Search for event lists and their \glspl{IRF} of CTAO South observations at energies above 10 TeV for blind search of PeVatrons from a data release using DataLink} {\em Identify all event lists and their associated \glspl{IRF} taken by CTAO South that contains events above 10 TeV. Data taken with the Small Size Telescopes or Medium Size Telescopes can be then selected. \/} +%proposal regular Obscore table joined to ivoa.obscore-hea \medskip \noindent Find all CTAO datasets satisfying: @@ -154,18 +216,17 @@ \subsubsection{Use Case --- Search for event lists and their \glspl{IRF} of CTAO The query output is a VOTable that follows the DataLink VO standard. We process this VOTABLE to access to the data: - \begin{enumerate}[(i)] - \item for each row of the query output, get the ``obs\_id'' and the ``access\_url'' of the DataLink describing the ObsCore dataset entry, + \item for each row of the query output, get the {\em obs\_id\/} and the {\em access\_url\/} of the DataLink describing the ObsCore dataset entry, \item get the DataLink VOTable showing the datasets associated to this entry - \item for each row of the DataLink VOTable, get the ``content\_qualifier'' and the ``access\_url'' column's value, - \item download the data associated to each ``access\_url'' value. + \item for each row of the DataLink VOTable, get the {\em content\_qualifier\/} and the {\em access\_url\/} attribute's value, + \item download the data associated to each {\em access\_url\/} value. \end{enumerate} Table \ref{tab:datalink1} displays an example of the DataLink response table attached to such an hea-event-list discovery. -The obs\_publisher\_did of the single discovered hea-event-list is repeated in the ID column of the DataLink table. -Mandatory FIELDS service\_def and error\_messsage are omitted because they are empty. +The {\em obs\_publisher\_did\/} of the single discovered {\bf hea-event-list} is repeated in the ID column of the DataLink table. +Mandatory FIELDS {\em service\_def\/} and {\em error\_messsage\/} are omitted as they are empty. \begin{landscape} \begin{center} @@ -190,9 +251,11 @@ \subsubsection{Use Case --- Search for event lists and their \glspl{IRF} of CTAO \end{center} \end{landscape} -\subsubsection{Use Case --- Search for spatially resolved spectropolarimetric observations of the Crab with spectral resolution R > 100} +\subsubsection{Use Case --- Search for spatially resolved spectro-polarimetric observations of the Crab with spectral resolution R > 100} -{\em Identify all event bundles for observations of the Crab that intersect the 1.0--100.0 keV energy range, have calibrated spatial and time axes, are spatially resolved in 2 dimensions in equatorial coordinates, have spectral resolution $R>100$, and include polarimetry measurements. Note that ObsCore specifies that the axes lengths --- s\_xel1, s\_xel2, em\_xel, t\_xel, pol\_xel --- should be set to $-1$ for non-pixelated data like event lists, so these quantities are not useful for this query.\/} +{\em Identify all event bundles for observations of the Crab that intersect the 1.0--100.0 keV energy range, have calibrated spatial and time axes, are spatially resolved in 2 dimensions in equatorial coordinates, have spectral resolution $R>100$, and include polarimetry measurements. Note that ObsCore specifies that the axes lengths --- {\em s\_xel1, s\_xel2, em\_xel, t\_xel, pol\_xel} --- should be set to $-1$ for non-pixelated data like event lists, so these quantities are not useful for this query.\/} + +%proposal regular Obscore table joined to ivoa.obscore-hea \medskip \noindent Find all datasets satisfying: @@ -225,6 +288,7 @@ \subsubsection{Use Case --- Search for spatially resolved spectropolarimetric ob \subsubsection{Use Case --- Identify PSF response-functions for further analysis of previously downloaded data products} {\em Identify all Chandra Source Catalog point spread functions for source detections that fall within 2 arcmin radius of (83.84358, $-5.43639$) in the Orion star-forming complex for Chandra observation 4374. These PSFs will be used to analyze previously downloaded catalog data products for the same field.\/ } +%proposal 2 : all in ivoa.obscore with 3 vocabularies but distinguish response-type by dataproduct\_subtype \medskip \noindent Find all datasets satisfying: @@ -250,7 +314,7 @@ \subsubsection{Use Case --- Identify PSF response-functions for further analysis \subsubsection{Use Case --- Get all the \glspl{IRF} for a given CTAO observation, for simulation purposes} -{\em Simulations are frequently used to estimate the science performance for a given astrophysical use case. To realise such simulations, \glspl{IRF} are required.\/ } +{\em Simulations are frequently used to estimate the science performance for a given astrophysical use case. To realize such simulations, \glspl{IRF} are required.\/ } \medskip \noindent Find the CTAO datasets satisfying: @@ -271,6 +335,7 @@ \subsubsection{Use Case --- Get all the \glspl{IRF} for a given CTAO observation AND (obs_id = '4374') AND (obs_collection = 'CTAO-DR1') \end{verbatim} +%can be implemented with the response table only \subsubsection{Use Case --- Search for all ANTARES neutrino data products for a given collection in the direction of a point source} @@ -306,7 +371,7 @@ \subsubsection{Use Case --- Retrieve the instrument response functions for a com \item dataproducts from the ``ARCA'' instrument of the ``KM3NeT'' facility, \item t\_min/t\_max from 2027--2030, ({\em i.e.\/}, MJD 61406--62870), \item event\_type = ``track'', - \item analysis mode optimised for pointsource searches. + \item analysis mode optimized for pointsource searches. \end{enumerate} \begin{verbatim} @@ -355,6 +420,7 @@ \subsubsection{Use Case --- Calculate the probability for a source class to be e {\em Using a catalog of potential sources, calculate the probability of measuring a $\nu_{\tau}$ neutrino flux from a stacking of all sources of that type with 10 years of data taking with widely spaced, high energy optical detectors like KM3NeT/ARCA.} +% proposal 2 \medskip \noindent Find all neutrino datasets satisfying: \begin{enumerate}[(i)] @@ -379,8 +445,8 @@ \subsection{Advanced Data Products} \subsubsection{Use Case --- Search for Chandra Source Catalog position error MCMC draws for X-ray detections in the vicinity of Gaia DR3 486718823701242368} -{\em Identify all Chandra Source Catalog position error MCMC draws for source detections that fall within 5 arcsec radius of (54.036061, $+61.907633$). The MCMC draws will be evaluated to establish whether there are potentially unresolved X-ray sources that may conincide with the white dwarf for observation planning.\/} - +{\em Identify all Chandra Source Catalog position error MCMC draws for source detections that fall within 5 arcsec radius of (54.036061, $+61.907633$). The MCMC draws will be evaluated to establish whether there are potentially unresolved X-ray sources that may coincide with the white dwarf for observation planning.\/} +% proposal 1 obscore only \medskip \noindent Find all datasets satisfying: \begin{enumerate}[(i)] @@ -399,14 +465,14 @@ \subsubsection{Use Case --- Search for Chandra Source Catalog position error MCM AND (dataproduct_subtype = 'poserr') AND (obs_collection = 'CSC2') \end{verbatim} -% mireille naive question: do we really need the join with the ivoa.obscore_hea table ? -% constraints on event_type would require it +% mireille: naive question: do we really need the join with the ivoa.obscore_hea table ? +% constraints on event_type would require it but no criteria from ivoa.obscore_hea is used \subsubsection{Use Case --- Search for flux maps for CTAO-North observations between two observations ID} -{\em Identify all flux maps from the CTAO-North data collection within a range of observation identifiers selected by the user .\/} - +{\em Identify all flux maps from the CTAO-North data collection within a range of observation identifiers selected by the user.\/} +% proposal obscore only would work with Obscore 1.1 \medskip \noindent Find all datasets satisfying: \begin{enumerate}[(i)] @@ -438,6 +504,7 @@ \subsubsection{Use Case --- Search for M31 source light curves and aperture phot %mireille %\TODO{when data product is light-curve, use the new term "light-curve" of the product-type vocabulary instead of "timeseries" which is the top concept in the vocabulary. This vocabulary should be adopted for the new implementations of ObsCore 1.1 and the next versions. } +% proposal 1 with join to ivoa.obscore-hea \medskip \noindent Find all datasets satisfying: \begin{enumerate}[(i)] @@ -451,7 +518,8 @@ \subsubsection{Use Case --- Search for M31 source light curves and aperture phot \begin{verbatim} -SELECT * FROM ivoa.obscore +SELECT obs_publisher_did, dataproduct_type, calib_level, energy_min, energy_max, +t_min, t_max, access_url FROM ivoa.obscore NATURAL JOIN ivoa.obscore_hea WHERE (CONTAINS(POINT(s_ra, s_dec), CIRCLE(10.6847, +41.2688, 1.5)) = 1) @@ -464,7 +532,7 @@ \subsubsection{Use Case --- Search for M31 source light curves and aperture phot \subsubsection{Use Case --- Search for the CTAO flux light curves of PKS 2155-304 in 2030} {\em Identify all light curves obtained on the source PKS 2155-304 in 2030 with the CTAO observatory.\/} - +% proposal ivoa.obscore with join to ivoa.obscore-hea \medskip \noindent Find all datasets satisfying: \begin{enumerate}[(i)]