Welcome to the SSCC!
If this is your first SSCC News, we welcome you to the Social Science Computing Core! Whether you’re a new faculty member, a new grad student, or an undergrad taking a class that uses SSCC services, we’re here to help you succeed. Visit the SSCC web site to learn about the services we provide.
SSCC News comes out roughly once every two months. If you’d rather not receive it, feel free to reply with ‘unsubscribe’. But if you’re not interested because you no longer need your SSCC account, let us know and we’ll close it for you.
SSCC Training
SSCC’s fall training is underway. A few highlights:
Data Visualization in R for Researchers had to be rescheduled to September 16th, so it’s not too late to register. This new workshop won’t just teach you R’s powerful ggplot package; it will teach you how to communicate your data and results effectively. It also covers plots for regression diagnostics and the ggeffects package. Python users who are interested in data visualization should also take this workshop, as you can implement what it teaches in plotnine (a Python port of ggplot).
Testing Data Wrangling Code in Stata will help you answer one of the hardest questions SSCC’s stat consultants get: “Is this right?” It’s also the SSCC’s first workshop specifically designed for an AI world. It’s not about AI, but we consider the practices it teaches to be essential if you’re going to let AI write research code for you. (Look for an R version of the workshop in the near future, but the focus is on practices, not code, and anyone who writes data wrangling code could benefit from attending.)
Caitlin Tefft will once again teach her popular series of workshops on Qualitative Data Analysis, including introductions to NVivo and MAXQDA. We are trying to make support for qualitative research a permanent part of the SSCC’s services, but it’s possible this will be the last time these workshops are offered. If you are interested in qualitative work, take advantage of Caitlin’s training and consulting now!
Data Impact Forum
The Data Impact Forum (September 21, 5:30-7:30) will showcase UW-Madison’s data consultancies & services, including SSCC, IRP, the Data Science Institute, the Statistical Consulting Group, and many more. Come learn about the data services that can impact your research!
Posit Assistant Now Supported by SSCC
RStudio on all SSCC servers now supports Posit Assistant, though you’ll need to pay for your own subscription to use it. Posit Assistant is essentially Claude with RStudio itself as the “agent harness” (the tool the AI uses to interact with the world). The SSCC’s statistical consultants have been using it and will do their best to answer questions about it.
As some of you already know, it’s possible to use other AI tools with SSCC servers. The SSCC’s statistical consultants cannot be experts on all the available tools, but feel free to send in questions and we’ll do our best. Even the questions we cannot answer are useful in helping us decide what to learn next.
Data Science Help Desk
The Data Science Institute’s Facilitation Team (formerly the Data Science Hub) has started offering Help Desk hours. Just drop in (no appointment needed):
Tuesdays, 2:30-4:30, Morgridge Hall Commons (Room 2564C)
Thursdays, 2:30-4:30, Memorial Library, Room 240
Whether you have a bug in your Python code, want to learn how to use git, or are curious about AI, DSI’s facilitators can get you pointed in the right direction.
The SSCC Help Desk is Now the ITS Service Desk
The L&S Information Technology Services Service Desk is now handling all questions and problem reports sent to the SSCC Help Desk. You can continue to use helpdesk@ssc.wisc.edu; in fact it’s helpful to them if questions about SSCC services go to that address. But feel free to use any of the contact methods for L&S ITS.
Note that there is no longer a drop-in help desk in 4226 Sewell Social Sciences Building, but the Service Desk has longer hours than the old SSCC Help Desk.
Record Slurm Usage, and Adapting To It
August set a new record for usage of the Slurm cluster, over 1.4 million core-hours. We’re happy to see the cluster put to work, but it has sometimes led to increased wait times. In the past we’ve nudged you towards using more Slurm resources rather than letting them sit idle, but with the cluster busier now it’s time to think more about efficiency, which will also reduce your wait time.
- Usually what prevents Slurm from running more jobs is not cores, but memory. Don’t reserve more memory than you actually need. Pay close attention to the efficiency report in the email you get when your job finishes: if your memory efficiency is less than about 90%, you can reserve less memory next time.
- For many tasks, adding more cores has diminishing marginal returns. If your Stata job only runs 20% faster with 64 cores rather than 32, reserve 32 when the cluster is busy–or if you have enough jobs to make it busy. Again, use the efficiency report: if your CPU efficiency is low, you can probably reserve fewer cores without much impact on performance.
- The
econpartition is a lower priority than eitherecon-gradorecon-fac. It allows you to use more servers, but if there are a lot of jobs inecon-grad, jobs submitted toeconwill frequently be preempted. - Jobs submitted to
econ-gradget high priority on Econ servers, but cannot use SSCC servers. When there are many jobs inecon-grad, jobs submitted to thessccpartition may run sooner. - Don’t forget the
shortpartition (maximum job length 6 hours). Jobs submitted toshortcan use any server, but there’s a server reserved for short jobs and sometimes it’s idle when the rest of the cluster is busy. - A task that’s been broken up into many small jobs can fill in the cracks between big jobs and take advantage of cores and memory that would otherwise sit idle.
- Use Slurm Status to submit jobs strategically, like sizing your jobs to fit what’s available. Run
squeue -aon Linstat (or LinSilo) to see which partitions are busy.