library(tidyverse)
library(sf)
library(geospaar)
districts <- read_sf(
  system.file("extdata", "districts.geojson", package = "geospaar")
)
farmers <- read_csv(
  system.file("extdata", "farmer_spatial.csv", package = "geospaar")
) %>% group_by(uuid) %>% 
  summarize(x = mean(x), y = mean(y), n = n()) %>%
  filter(y > -18) #%>% st_as_sf(coords = c("x", "y"), crs = 4326)
p <- ggplot() + 
  geom_sf(data = districts, lwd = 0.1) + 
  geom_point(data = farmers, 
             aes(x = x, y = y, size = n * 0.8, color = n), alpha = 0.9) +
  scale_color_viridis_c(guide = FALSE) + theme_void() + 
  theme(legend.position = c(0.85, 0.2)) +
  scale_size(range = c(0.1, 5), name = "N reports/week")
ggsave(here::here("docs/figures/zambia_farmer_repsperweek.png"), 
       width = 6, height = 4, dpi = 300, bg = "transparent")

Today

  • Assignment review
  • Homework results
  • Continuing on control structures with emphasis on *apply

Show your homework

set.seed(10)
g <- ...

Data generation

Create the following:

  • dat, a data.frame built from V1, V2, V3, and V4, where:
    • V1 = 1:20
    • V2 is a random sample between 1:100
    • V3 is drawn from a random uniform distribution between 0 and 50
    • V4 is a random selection of the letters A-E
    • Use set.seed(50)
  • Do this all at once (i.e. wrap the creation of V1-V4 in the data.frame call, precede it with set.seed())

Check the answer

set.seed(10)
dat <= dat.frame(
  V1 = 1;20, 
  V2 = sample(1:100, size = 20, replace = True),
  V3 = runif(n = 20, min = 0, max = 50), 
  V4 = sample(letters(1:5), size = 20)
)

Advanced

  • Use lapply to make three data.frames captured in a list l, each composed of one randomly sampled column v1 (selecting from integers 1:10, with length = 20), and the second being v2 composed of lowercase letters, randomly selected using sample, also of length 20.
  • The iterator in the lapply should be 10, 20, 30, which become the random seeds for the sampling (in the body of the lapply)

Check the code

l <- lapply(10:30, function(x) {
  data.frame(
    V1 = sample(1:10, length = 20), 
    v2 = sample(letters, size = 20, Replace = TRUE)
  )
}}

Exercises

  • Use a for to iterate over each row of dat and calculate it’s sum
  • Do the same with lapply and sapply
  • Do the same using rowSums
  • Select rows from dat containing the letter “E” in V4, and take the mean of values from the result in column V3
  • Create a function called myfun (just in your script, not as a package function). Have it add 20% of x (input value) to x. Apply it to all eligible values in dat