library(sf)
library(dplyr)
library(ggplot2)
library(rnaturalearth)
library(rnaturalearthdata)
data(world.cities, package = "maps")

world <- ne_countries(scale = "medium", returnclass = "sf")
afr_capitals <- world.cities %>% filter(capital == 1) %>% 
  st_as_sf(coords = c("long", "lat"), crs = 4326) %>% 
  st_intersection(., world %>% filter(continent == "Africa"))
p <- world %>% filter(continent == "Africa") %>% 
  ggplot() + geom_sf(aes(fill = name), lwd = 0.2) + 
  geom_sf(data = afr_capitals, col = "blue", size = 0.5) + 
  scale_fill_grey(guide = FALSE) + theme_minimal()
ggsave(here::here("external/slides/figures/africa_capitals.png"), 
       width = 5, height = 4, dpi = 300, bg = "transparent")

Today

  • Indexing and subsetting
  • Control structures

Create data structures

  • Starting note: Use set.seed(10) when creating random data
  • Create a list l from a, b, d.
    • a: a random vector of integers with 10 elements drawn from 1-20, with vectors named “V1” through “V10”:
    • b: Using a as an index to select from letters
    • d: Use rnorm with a mean = 100 and an sd of 20
    • Assign the names of the vectors in l to the l’s elements
  • m: a matrix with three integer columns named “V1”, “V2”, “V3”
    • Create each column first as its own vector, then combine
    • V1 = 1:10
    • V2 is a random sample between 1:100
    • V3 is drawn from a random uniform distribution between 0 and 50
  • dat, a data.frame built from V1, V2, V3, and V4
    • V4 is a random selection of the letters A-E

Functions

Components

function_name <- function(arg1, arg2 = 1:10, 
                          arg3 = ifelse(arg2 == 2, TRUE, FALSE)) {
  body
}

Three components of a function:

  • formals(): arguments
  • body(), the code, which returns the last object generated, unless specified with return(x).
  • environment(), function finds the values

Unnamed functions are anonymous functions. (Used in *apply)

Using x in a function does not change its global value.

x <- 1:10
myfun <- function() {
  x * 10
}
myfun()
 [1]  10  20  30  40  50  60  70  80  90 100
myfun <- function(x) {
  x <- x * 10
  return(x)
}
x <- 10
myfun(x = 20)
[1] 200
x
[1] 10

Each time you run myfun, a new function environment is created.

myfun <- function(x) {
  x <- x * 10
  print(environment())
  return(x)
}
myfun(x)
<environment: 0x1176ca5c0>
[1] 100
myfun(x)
<environment: 0x1176fff58>
[1] 100

Global assignment.

Use <<- to change value of global variable within a function.

a <- 10
myfun <- function(x) {
  a <<- x * 10   ## note <<- instead of <- 
  return(a)
}
myfun(5)
[1] 50
print(a)
[1] 50

Useful functions

which finds indices where a condition is true.

v <- 10:15
print(v)
[1] 10 11 12 13 14 15
a <- which(v %% 3 == 0) ## subset to elements divisible by 3
print(a) ## shows indices where condition is true.
[1] 3 6

Useful functions

which.min finds index of min value

v <- sample(1:20, 10)
print(v)
 [1]  4 12 16  8 17  6 13 14  7 18
print(which.min(v)) # index of min value
[1] 1
print(which.max(v)) # index of max value
[1] 10

data.frame vs data.table vs. tibble

  • all 2D structures.
  • data.frame = Base R
  • tibble = tidyverse
  • data.table = very fast

For now, we’ll stick to data.frame

data.frame indexing

data.frame uses the following to subset: [*row conditions, *column conditions]

df <- data.frame(v1 = 1:5, v2 = 6:10)
rownames(df) <- LETTERS[1:5]
print(df)
  v1 v2
A  1  6
B  2  7
C  3  8
D  4  9
E  5 10

data.frame indexing

  • Index using names.
  • Empty index [ , 'v2] means “keep all rows”
df[,'v2'] ## column indexing
[1]  6  7  8  9 10
df[c("A", "B", "D"), ] ## row indexing
  v1 v2
A  1  6
B  2  7
D  4  9

data.frame subset

  • Logical subset
df[df$v1 > 3, ] ## get observations (rows) where first column is larger than 3
  v1 v2
D  4  9
E  5 10

Control structures

Branching

  • Pay attention to { } placement
a <- 5
if(a > 10) {
  print("Greater than 10!")
} else {
  print("Less than or equal to 10")
}
[1] "Less than or equal to 10"

Looping

b <- 1:3
for(i in b) print(i)
[1] 1
[1] 2
[1] 3
b <- 1:5
a <- 2
for(i in b){
  a <- 2 * a
  print(a)
}
[1] 4
[1] 8
[1] 16
[1] 32
[1] 64

*apply

  • A special form of looping
  • Intended for applying a function to data. Uses anonymous function.
  • 3 main kinds: sapply, lapply, apply

sapply

sapply iterates over input and returns a vector.

v <- 1:10
sapply(v, function(x) x + 10) ## adds 10 to each element in v.
 [1] 11 12 13 14 15 16 17 18 19 20

Use { } for more complicated functions. BUT be careful with order of { }, ( )

v1 <- 1:10
v2 <- sapply(v1, function(x){
  y <- x^2 
  return(y)
}) #
print(v2)
 [1]   1   4   9  16  25  36  49  64  81 100

sapply

If you don’t specify return, the last object created will be returned.

v1 <- 1:10
v2 <- sapply(v1, function(x){
  y <- x^2  ## y will be returned
}) #
print(v2)
 [1]   1   4   9  16  25  36  49  64  81 100

lapply

  • Similar to sapply, except final object is returned as list.
  • Useful if you need to store more complex objects (data.frame, plot, raster etc.)
v1 <- 1:10
v2 <- lapply(v1, function(x){
  y <- x^2  ## y will be returned
}) #
print(v2)
[[1]]
[1] 1

[[2]]
[1] 4

[[3]]
[1] 9

[[4]]
[1] 16

[[5]]
[1] 25

[[6]]
[1] 36

[[7]]
[1] 49

[[8]]
[1] 64

[[9]]
[1] 81

[[10]]
[1] 100

apply

apply works well for 2D data, when you want to apply function over a row or column.

v1 <- sample(1:100, 10)
v2 <- sample(1:100, 10)
DF <- data.frame(v1, v2) ## data frame columns will take names of vectors
DF
   v1 v2
1  48 13
2  30 18
3   8 34
4   4 78
5   3 82
6  59 33
7  97 31
8  36 57
9  56 74
10 49 45

Use apply to get column max value. The index 2 means “apply function to columns”.

colMax <- apply(DF, 2, FUN = max)
colMax
v1 v2 
97 82 

Use apply to get row max value. The index 1 means “apply function to rows”.

rowMax <- apply(DF, 1, FUN = max)
rowMax
 [1] 48 30 34 78 82 59 97 57 74 49

We can use apply or sapply to create a new column in a data frame.

DF$rowMax <- apply(DF, 1, FUN = max)
DF
   v1 v2 rowMax
1  48 13     48
2  30 18     30
3   8 34     34
4   4 78     78
5   3 82     82
6  59 33     59
7  97 31     97
8  36 57     57
9  56 74     74
10 49 45     49

Class exercises (work in threes)

Create data

  • Use a random seed of 10 for generating data
  • Create a vector g with a random sample of 10 numbers drawn from 1-5.
  • Create a data.frame called dat combining g with h, where h = g divided by a number drawn using rnorm with a mean of 10 and sd of 1.
  • Create a list l2 that combines a and dat, naming each list element by its object name.

Use the data as follows:

  • Write a for loop that iterates through g and prints only the values >2, otherwise it should print “Nope!”
  • Use apply to find the maximum value of each row in dat
  • Use sapply to find the maximum value of each row in dat
  • Use lapply to iterate through each element of l2 and subtract the minimum value of each element from the maximum value of each element