Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
27 commits
Select commit Hold shift + click to select a range
d899bd4
added the birdnet_add_site() function
SunnyTseng Jun 12, 2026
4943d8a
added roxygen for documentation
SunnyTseng Jun 12, 2026
3353579
fine tune the notes
SunnyTseng Jun 12, 2026
49c2b99
Add the new function in the documentation
SunnyTseng Jun 12, 2026
e247482
Added a new function birdnet_get_effort (export)
SunnyTseng Jun 12, 2026
fc818dc
birdnet_get_effort documentation, name space
SunnyTseng Jun 13, 2026
844e732
initial code for the detection history from template to the official …
SunnyTseng Jun 13, 2026
3eba74d
initial check throughout the code
SunnyTseng Jun 13, 2026
3c6b77e
align the code
SunnyTseng Jun 15, 2026
eda744d
finalized the code content
SunnyTseng Jun 15, 2026
c210193
updated the detection history documentation
SunnyTseng Jun 15, 2026
4f007e5
Update the documentation
SunnyTseng Jun 15, 2026
e6563ad
update parameter definition
SunnyTseng Jun 15, 2026
05f45ae
Updated dependency
SunnyTseng Jun 15, 2026
8513d50
Added effort data as example datasets
SunnyTseng Jun 15, 2026
8fd381a
Added data documentation
SunnyTseng Jun 15, 2026
1293158
reduced the species number in the example dataset
SunnyTseng Jun 16, 2026
e0e3164
format the function
SunnyTseng Jun 18, 2026
13e8267
added effort matrix as one additional output from detection_matrix() …
SunnyTseng Jun 19, 2026
67b4769
Adopt the min_unique_days
SunnyTseng Jun 19, 2026
0fee008
Updated the function documentation
SunnyTseng Jun 19, 2026
de1bf37
Added test functions for birdnet_add_site()
SunnyTseng Jun 25, 2026
1915daf
cteated test file for detection_history() function
SunnyTseng Jun 25, 2026
127fcbe
Added checkmate argument check for detection history
SunnyTseng Jun 25, 2026
ffd89f5
initial test functions for detection_history
SunnyTseng Jun 25, 2026
5adbb04
updated argument check and the test functions for get_effort()
SunnyTseng Jun 25, 2026
1e2b04f
added withr as suggested dependency
SunnyTseng Jun 25, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions DESCRIPTION
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,6 @@ Description: What the package does (one paragraph).
License: MIT + file LICENSE
Encoding: UTF-8
Roxygen: list(markdown = TRUE)
RoxygenNote: 7.3.2
URL: https://birdnet-team.github.io/birdnetTools/, https://github.com/birdnet-team/birdnetTools
BugReports: https://github.com/birdnet-team/birdnetTools/issues
Imports:
Expand All @@ -28,13 +27,16 @@ Imports:
shinyFiles,
shinyWidgets,
stringr,
tidyr,
tuneR
Suggests:
knitr,
rmarkdown,
testthat (>= 3.0.0)
testthat (>= 3.0.0),
withr
Config/testthat/edition: 3
VignetteBuilder: knitr
Depends:
R (>= 4.1.0)
LazyData: true
Config/roxygen2/version: 8.0.0
3 changes: 3 additions & 0 deletions NAMESPACE
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,9 @@
export(birdnet_add_datetime)
export(birdnet_calc_threshold)
export(birdnet_combine)
export(birdnet_detection_history)
export(birdnet_filter)
export(birdnet_get_effort)
export(birdnet_heatmap)
export(birdnet_launch_validation)
export(birdnet_subsample)
Expand All @@ -18,6 +20,7 @@ importFrom(bslib,sidebar)
importFrom(cli,cli_alert_success)
importFrom(cli,cli_alert_warning)
importFrom(cli,cli_li)
importFrom(dplyr,.data)
importFrom(dplyr,bind_rows)
importFrom(dplyr,mutate)
importFrom(dplyr,select)
Expand Down
222 changes: 222 additions & 0 deletions R/birdnet_detection_history.R
Original file line number Diff line number Diff line change
@@ -0,0 +1,222 @@
#' Generate Detection History Matrix, Effort Matrix, and Summary for Occupancy Modeling
#'
#' Summarizes BirdNET detection data across specified survey intervals (occasions),
#' filters sites based on minimal detection persistence thresholds, and aligns them
#' with operational effort data. Returns a zero-filled site-by-occasion binary matrix,
#' an identical matching matrix documenting sampling effort intensity for modeling
#' detection probability covariates, and a detailed long-format data frame summary.
#'
#' @details
#' The function groups continuous temporal data into distinct survey blocks using
#' `lubridate::floor_date()`. Detections are cross-referenced against your
#' `effort_data`: occasions where monitoring effort occurred but no target
#' species were detected are explicitly zero-filled. If an ARU was not operational
#' during a specific time block, it is preserved as an `NA` value in the detection
#' history to ensure structural integrity for missing-visit designs.
#'
#' Values greater than 0 in the final detection matrix are collapsed to `1` to format
#' the output for binary presence/absence occupancy models (e.g., `spOccupancy`, `unmarked`).
#'
#' @param data A data frame containing BirdNET detections, including column matches
#' for filepaths and prediction confidence scores.
#' @param effort_data A data frame containing monitoring operational effort,
#' requiring at least `site` and `date` columns to indicate the active
#' monitoring windows and locations of each ARU device. Users can generate
#' this via [birdnet_get_effort()], which derives effort data from a directory
#' of audio files by defining a site-date combination as "active" if at least
#' one recording exists. If an `n_files` column is present, file counts will
#' be aggregated per survey occasion block.
#' @param survey_interval A character string specifying the temporal unit for
#' grouping survey occasions (e.g., `"1 day"`, `"1 week"`, `"7 days"`).
#' Passed directly to \code{\link[lubridate:floor_date]{lubridate::floor_date()}}.
#' @param i An integer specifying the path hierarchy index for extracting site IDs.
#' Passed directly to \code{\link{birdnet_add_site}}. Defaults to `-2`.
#' @param min_unique_days An integer specifying the threshold of unique calendar days
#' a site must possess raw detections on to be kept. Sites with detections spanning fewer
#' than `min_unique_days` are dropped early from compilation. Defaults to `1`.
#'
#' @return A named `list` containing three components:
#' \describe{
#' \item{detection_history}{A numeric base R `matrix` where rows represent
#' unique sites (assigned as row names), columns represent chronological temporal
#' occasions, and cells indicate binary occupancy integers (`1`, `0`,
#' or `NA` for missing effort).}
#' \item{effort_matrix}{A numeric base R `matrix` matching the exact dimensions and
#' sorting order of `detection_history`. If `n_files` was present in the effort data,
#' cells represent total file counts per site-occasion. Otherwise, cells contain binary
#' integers indicating presence (`1`) or absence (`0`) of operational effort.}
#' \item{detection_summary}{A data frame in long format containing the underlying
#' aggregated metrics per site/occasion, including detection counts (`n_detections`),
#' maximum verification confidence (`max_conf`), and the file path of the
#' highest confidence detection (`max_conf_audio`).}
#' }
#'
#' @importFrom dplyr .data
#' @export
birdnet_detection_history <- function(data,
effort_data,
survey_interval,
i = -2,
min_unique_days = 1) {


# argument check ----------------------------------------------------------

# 1. Check data is a data frame with required columns
checkmate::assert_data_frame(data)

cols <- birdnet_detect_columns(data)
required_cols <- c("confidence", "filepath")
missing_cols <- required_cols[is.na(cols[required_cols])]

if (length(missing_cols) > 0) {
rlang::abort(
paste0(
"The input data is missing required BirdNET columns: ",
paste(missing_cols, collapse = ", "),
". Please provide a valid BirdNET output data frame."
)
)
}


# 2. Check effort_data is a data frame with required columns
checkmate::assert_data_frame(effort_data)

effort_cols <- colnames(effort_data)
required_effort_cols <- c("site", "date")
missing_effort_cols <- setdiff(required_effort_cols, effort_cols)

if (length(missing_effort_cols) > 0) {
rlang::abort(
paste0(
"The input effort data is missing required columns: ",
paste(missing_effort_cols, collapse = ", "),
". Please provide a valid effort data frame."
)
)
}

# 3. Check survey_interval is a character string following lubridate units
checkmate::assert_string(survey_interval, min.chars = 1)
if (!stringr::str_detect(survey_interval, "^\\d*\\s*(day|week|month|year|hour|minute)s?$")) {
rlang::abort(
paste0(
"`survey_interval` must be a valid lubridate unit string (e.g., '1 day', '2 weeks'). ",
"You provided: '", survey_interval, "'."
)
)
}

# 4. Check i is an integer
checkmate::assert_int(i, tol = 0)

# 5. Check min_unique_days is a positive integer
checkmate::assert_int(min_unique_days, lower = 1, tol = 0)





# main function -----------------------------------------------------------

cols <- birdnet_detect_columns(data)

# 1. Summarize detections by site and occasion block
detections_summarized <- data |>
birdnet_add_site(i = i) |>
birdnet_add_datetime() |>
# filter to only include sites with detections from more than n days
dplyr::group_by(.data$site) |>
dplyr::filter(dplyr::n_distinct(.data$date) >= min_unique_days) |>
dplyr::ungroup() |>
# group detections into survey occasions based on the specified interval
dplyr::mutate(occasion = lubridate::floor_date(x = .data$date,
unit = survey_interval)) |>
dplyr::group_by(.data$site, .data$occasion) |>
dplyr::summarise(n_detections = dplyr::n(),
max_conf = max(.data[[cols$confidence]], na.rm = TRUE),
max_conf_audio = .data[[cols$filepath]][which.max(.data[[cols$confidence]])],
.groups = "drop")



# 2. Process effort
baseline_effort <- effort_data |>
dplyr::mutate(occasion = lubridate::floor_date(x = .data$date,
unit = survey_interval))
if ("n_files" %in% names(baseline_effort)) {
# if n_files exists, aggregate the total file counts per site/occasion
baseline_effort <- baseline_effort |>
dplyr::group_by(.data$site, .data$occasion) |>
dplyr::summarise(n_files = sum(.data$n_files, na.rm = TRUE), .groups = "drop")
} else {
# if n_files is missing, simply keep unique combinations of site and occasion
baseline_effort <- baseline_effort |>
dplyr::distinct(.data$site, .data$occasion)
}



# 3. join detections, fill zeros, and pivot wide
detections_zero_filled <- baseline_effort |>
# left join ensures we only evaluate occasions where the devices were running
dplyr::left_join(detections_summarized, by = c("site", "occasion")) |>

# differentiate true zeros from missing effort
dplyr::mutate(n_detections = tidyr::replace_na(.data$n_detections, 0),
max_conf = tidyr::replace_na(.data$max_conf, 0),
max_conf_audio = tidyr::replace_na(.data$max_conf_audio, "none"))



# 4. Creating matrix structure: pivot to wide format with sites as rows and occasions as columns

# isolate matrix structure and shape wide for modeling packages (e.g., unmarked and spOccupancy)
detection_history_df <- detections_zero_filled |>
dplyr::select("site", "occasion", "n_detections") |>
# manipulate n_detections column to make it 1 if it's larger than 0,
# otherwise 0 (for occupancy modeling)
dplyr::mutate(n_detections = ifelse(.data$n_detections > 0, 1, 0)) |>
dplyr::arrange(.data$occasion, .data$site) |>
tidyr::pivot_wider(id_cols = "site",
names_from = "occasion",
values_from = "n_detections",
values_fill = NA)

detection_history <- as.matrix(detection_history_df[, -1])
rownames(detection_history) <- detection_history_df$site



# 5. Create the effort matrix for detection probability purpose
baseline_effort <- baseline_effort |>
dplyr::arrange(.data$occasion, .data$site)

if ("n_files" %in% names(baseline_effort)) {
# if n_files exists, we can use the file counts as a measure of effort
wide_effort <- baseline_effort |>
tidyr::pivot_wider(id_cols = "site",
names_from = "occasion",
values_from = "n_files",
values_fill = 0)
} else {
# if n_files is missing, we can only indicate presence (1) or absence (0) of effort
wide_effort <- baseline_effort |>
tidyr::pivot_wider(id_cols = "site",
names_from = "occasion",
values_from = "occasion",
values_fn = \(x) 1,
values_fill = 0)
}

# Strip the site column to create a clean matrix, keeping site names as rownames
effort_matrix <- as.matrix(wide_effort[, -1])
rownames(effort_matrix) <- wide_effort$site


return(list("detection_history" = detection_history,
"effort_matrix" = effort_matrix,
"detection_summary" = detections_zero_filled))
}

68 changes: 68 additions & 0 deletions R/birdnet_get_effort.R
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
#' Calculate recording effort by site and date
#'
#' Scans a directory for all audio files, extracts site and datetime
#' metadata from their paths/filenames, and returns a unique timeline
#' of recording effort with file counts.
#'
#' The function identifies audio files matching common extensions, automatically
#' detects site names via [birdnet_add_site()], parses dates via
#' [birdnet_add_datetime()], and reduces the output to a unique combination
#' of sites and dates, summarizing total files recorded.
#'
#' @param path A character string specifying the path to the directory
#' containing the audio files.
#' @param i An integer specifying the index of the path element to extract
#' as the site identifier when split by slashes. Defaults to `-2`, which
#' corresponds to the immediate parent directory of the file, passed directly
#' to [birdnet_add_site()]. Negative values count from the right-hand side.
#'
#' @return A tibble (data frame) with three columns:
#' \describe{
#' \item{site}{The extracted site identifier.}
#' \item{date}{The date on which recording effort occurred.}
#' \item{n_files}{Integer. The total number of audio files recorded at that
#' site on that specific date.}
#' }
#'
#' @examples
#' \dontrun{
#' effort_df <- birdnet_get_effort("path/to/audio/storage", i = -2)
#' head(effort_df)
#' }
#'
#' @export
birdnet_get_effort <- function(path, i = -2) {


# argument check ----------------------------------------------------------

# 1. Check path is a single, valid directory path string
checkmate::assert_string(path, min.chars = 1)
checkmate::assert_directory_exists(path, access = "r")


# 2. Check i is an integer
checkmate::assert_int(i, tol = 0)


# main function -----------------------------------------------------------

# find files and build the effort dataframe
effort <- path |>
# list the file names
list.files(recursive = TRUE,
full.names = TRUE,
pattern = "\\.(wav|mp3|m4a|flac|ogg|wma)$",
ignore.case = TRUE) |>
# convert to tibble for processing, extract time and location
(\(x) dplyr::as_tibble(data.frame(filepath = x, stringsAsFactors = FALSE)))() |>
birdnet_add_datetime() |>
birdnet_add_site(i = i) |>

# keep only the relevant columns and unique rows
dplyr::summarise(n_files = dplyr::n(), .by = c("site", "date"))


return(effort)
}

31 changes: 31 additions & 0 deletions R/data_documentation.R
Original file line number Diff line number Diff line change
Expand Up @@ -32,3 +32,34 @@
#' @source <https://sunnytseng.ca/>
"example_jprf_2023"



#' Example monitoring effort table from John Prince Research Forest
#'
#' A sample operational effort table mapping the active recording history of
#' Autonomous Recording Units (ARUs) deployed across 5 sites in John Prince
#' Research Forest, British Columbia, Canada, during May–June 2023. This data
#' documents the baseline monitoring effort, where a given location and date
#' combination is associated with an active ARU device if one or more audio files
#' were successfully recorded.
#'
#' This dataset acts as the operational counterpart to `example_jprf_2023` and is
#' useful for demonstrating workflow alignment between species detections and true
#' field effort, specifically for zero-filling non-detections in occupancy modeling.
#'
#' @details
#' This dataset was generated directly using the [birdnet_get_effort()] function.
#' For more details on the generation parameters, data constraints, and internal
#' file processing pipelines, please refer to the function documentation.
#'
#' @format ## `effort_jprf_2023`
#' A data frame with rows and columns detailing active recording days. Key columns include:
#' \describe{
#' \item{site}{Character string indicating the unique identifier for the ARU deployment location}
#' \item{date}{Date object representing the calendar day of monitoring effort}
#' \item{n_files}{Integer representing the total number of audio files recorded at that location on that day}
#' }
#'
#' @source <https://sunnytseng.ca/>
"effort_jprf_2023"

Loading
Loading