Skip to main content

Sports Brands and Eco-Sustainability on Twitter

Hakan Ates
Author
Hakan Ates
I build full-stack applications and LLM-powered tools.
A research project for the Network Science course at the University of Padua: do sports brands genuinely spread eco-sustainable values on Twitter, or is it greenwashing?

Highlights
#

  • Collected 110,580 English-language tweets with snscrape, matched against a hand-built list of 72 eco-sustainability hashtags.
  • Modeled brand perception as a weighted brand-hashtag co-occurrence network and analyzed it with NetworkX and Gephi: centrality, robustness under targeted node removal, and modularity-based community detection.
  • Located each brand’s real position in the semantic network: Patagonia and The North Face sit fully inside the eco-sustainability cluster, Nike only partially, Adidas in the undifferentiated core.

How it works
#

flowchart TD
    A["snscrape + 72 eco-hashtags"] --> B["110,580 English tweets"]
    B --> C["Human-coded cleaning"]
    C --> D["Brand-hashtag co-occurrence network
Full and Light projections"] D --> E["NetworkX: centrality, PageRank,
assortativity, robustness"] D --> F["Gephi: modularity communities,
Force Atlas 2"] B --> G["LIWC sentiment analysis"] E --> H["Findings"] F --> H G --> H

Case Study
#

Problem. Sports brands run loud eco-sustainability campaigns, but campaign volume says nothing about whether users actually connect a brand to green values. Our 8-person team, mixing social-science and technical backgrounds, asked that question directly on Twitter. Perception was read from the tweets themselves, not from the brands’ messaging.

Approach. I collected the tweet corpus with snscrape, covering everything from the first tweet on Twitter to the collection date, and ran the network analysis; I also contributed to the research design and writing. The 72 eco-sustainability hashtags were derived from the brands’ own campaign pages. We built a network with brands and hashtags as nodes and co-occurrence strength as edge weights, studying brands across categories (Nike, Adidas, Vans, plus Burton, Patagonia, The North Face, Quiksilver, Hurley, Oakley, Reef, Columbia, Dainese and others). The full graph was too large for our hardware, so we worked on two edge-weight-filtered projections: a Full network (weight > 1) and a Light network (weight > 15). We also split the dataset at 20 August 2018, the first Fridays for Future strike, to compare perception before and after.

What was hard. Three brand names (Omen, Reef, Columbia) are ordinary English words, so the scrape pulled in large amounts of semantically unrelated tweets. We only found this by going through tweets with human coding, then cleaning the dataset. The raw node count also exceeded what our machines could handle, which is what forced the filtered projections in the first place.

Outcome. The network is disassortative: hub nodes (brand names and top hashtags) connect to peripheral nodes rather than to each other, and the degree-distribution exponent rises from 2.15 in the full network to 3.01 in the green community, which therefore isn’t properly scale-free. Community detection via modularity and Force Atlas 2 found 5 clusters, confirmed by conductance, separating brands sharply by position. LIWC sentiment showed positive-emotion language far outweighing negative across brands, with Quiksilver, Hurley, Patagonia and Oakley above 3.5 on the posemo scale.

Links. Project report (PDF, Google Drive)

Stack
#

Python (snscrape, NetworkX), Gephi (modularity, Force Atlas 2), LIWC.